← Search

Qi Sun

34 accepted papers

2026

Ani3DHuman: Photorealistic 3D Human Animation with Self-guided Stochastic Sampling

CVPR 2026

Current 3D human animation methods fail at photorealism: kinematics-based approaches lack non-rigid dynamics like clothing, while methods reconstructing from generated videos suffer from low-quality artifacts and identity loss. To overcome these limitations, we present Ani3DHuman, a framework that m

Cited by 0SourcecodeScholar
2026

Circuit-Think: A Multimodal Reasoning Framework for Automated Circuit-to-Netlist Translation with Trajectory-Guided Reinforcement Learning

AAAI 2026technical

Vision Language Models (VLMs) have shown strong performance in multimodal understanding, offering promise for the circuit-to-netlist translation task. However, the diverse component symbols and complex connections in circuit images challenge VLMs in understanding physical layouts and reasoning for e

Cited by 0SourcePDFScholar
2026

Learning to Orchestrate Agents in Natural Language with the Conductor

ICLR 2026poster

Powerful large language models (LLMs) from different providers have been expensively trained and finetuned to specialize across varying domains. In this work, we introduce a new kind of Conductor model trained with reinforcement learning to automatically discover powerful coordination strategies amo…

Cited by 0SourceScholar
2026

LithoDreamer: A Physics-Informed World Model for Multi-Stage Computational Lithography

ICML 2026poster

As semiconductor technology nodes continue to shrink, computational lithography has become critical to yield and performance. However, real-world lithography is a continuous, multi-stage physical process driven by implicit interventions, which cannot be captured by the existing static or stage-wise …

Cited by 0SourceScholar
2026

Narrowing the ANN–SNN Gap for 1D Signal Classification with Multi-Scale Temporal Encoding and Sparsity-Regularized Transform Encoding

ICML 2026poster

Spiking neural networks (SNNs) promise energy-efficient inference, yet on static vision benchmarks they often trail matched ANNs under short simulation horizons. Under a matched-backbone and matched-budget protocol without extra tricks, we find that this ANN-SNN accuracy gap is consistently smaller …

Cited by 0SourceScholar
2026

Optimization Method for Surrogate Function in Spiking Neural Networks Based on Membrane Potential Distribution

AAAI 2026technical

Spiking Neural Networks (SNNs) offer promising energy efficiency and temporal sparsity for edge intelligence, but their training remains difficult due to gradient mismatch, membrane potential drift, and discretization errors. In this paper, we propose a membrane potential-guided surrogate optimizati

Cited by 0SourcePDFScholar
2026

SafeCompass: Dynamic Chain-of-Thought Steering via Inference-Time Safety Signals

ICML 2026poster

Large reasoning models (LRMs) achieve strong performance by explicitly generating chain-of-thought (CoT) reasoning, but this reasoning process can be manipulated by adversarial prompts. Inference-time CoT interventions offer a simple and lightweight approach to improving safety, yet existing methods…

Cited by 0SourceScholar
2026

ShoppingBench: A Real-World Intent-Grounded Shopping Benchmark for LLM-based Agents

AAAI 2026technical

Existing benchmarks in e-commerce primarily focus on basic user intents, such as finding or purchasing products. However, real-world users often pursue more complex goals, such as applying vouchers, managing budgets, and finding multi-products seller. To bridge this gap, we propose ShoppingBench, a

Cited by 0SourcePDFScholar
2025

Advancing Sequential Numerical Prediction in Autoregressive Models

ACL 2025short

Autoregressive models have become the de facto choice for sequence generation tasks, but standard approaches treat digits as independent tokens and apply cross-entropy loss, overlooking the coherent structure of numerical sequences. This paper introduces Numerical Token Integrity Loss(NTIL) to addre…

2025

Aurora-M: Open Source Continual Pre-training for Multilingual Language and Code

COLING 2025industry

Pretrained language models are integral part of AI applications, but their high computational cost for training limits accessibility. Initiatives such as Bloom and StarCoder aim to democratize access to pretrained models for collaborative community development. Despite these efforts, such models enc…

Cited by 2SourcePDFScholar
2025

Complete Coverage Path Planning Algorithm Based on Improved Biologically Inspired Neural Networks in Spray Painting

RA-L 2025

Intelligent putty coating technology is the main way to improve the degree of automation of railroad vehicle painting workshops. The two-component putty, which is currently used in the vehicle coating system, has extremely low fluidity, which requires improving the full coverage of the spray path wh

Cited by 3SourceScholar
2025

Concurrent Reinforcement Learning with Aggregated States via Randomized Least Squares Value Iteration

ICML 2025poster

Designing learning agents that explore efficiently in a complex environment has been widely recognized as a fundamental challenge in reinforcement learning. While a number of works have demonstrated the effectiveness of techniques based on randomized value functions on a single agent, it remains un…

Cited by 0SourcePDFScholar
2025

EG4D: Explicit Generation of 4D Object without Score Distillation

ICLR 2025poster

In recent years, the increasing demand for dynamic 3D assets in design and gaming applications has given rise to powerful generative pipelines capable of synthesizing high-quality 4D objects. Previous methods generally rely on score distillation sampling (SDS) algorithm to infer the unseen views a…

2025

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning

ACL 2025long

Traditional reinforcement learning-based robotic control methods are often task-specific and fail to generalize across diverse environments or unseen objects and instructions. Visual Language Models (VLMs) demonstrate strong scene understanding and planning capabilities but lack the ability to gener…

2025

From Grounding to Manipulation: Case Studies of Foundation Model Integration in Embodied Robotic Systems

EMNLP 2025

Foundation models (FMs) are increasingly applied to bridge language and action in embodied agents, yet the operational characteristics of different integration strategies remain under-explored—especially for complex instruction following and versatile action generation in changing environments. We i

2024

BiE: Bi-Exponent Block Floating-Point for Large Language Models Quantization

ICML 2024poster

Nowadays, Large Language Models (LLMs) mostly possess billions of parameters, bringing significant challenges to hardware platforms. Although quantization is an efficient approach to reduce computation and memory overhead for inference optimization, we stress the challenge that mainstream low-bit qu…

Cited by 5SourcePDFScholar
2024

RSAP-DFM: Regime-Shifting Adaptive Posterior Dynamic Factor Model for Stock Returns Prediction

IJCAI 2024poster

As the latest development of asset pricing research, how to use machine learning to improve the performance of factor models has become a topic of concern in recent years. The variability of the instantaneous macro environment brings great difficulties to quantitative investment, so the extended fac…

Cited by 2SourcePDFScholar
2024

Rocket Landing Control with Random Annealing Jump Start Reinforcement Learning

IROS 2024

Rocket recycling is a crucial pursuit in aerospace technology, aimed at reducing costs and environmental impact in space exploration. The primary focus centers on rocket landing control, involving the guidance of a nonlinear under-actuated rocket with limited fuel in real-time. This challenging task

Cited by 6SourceScholar
2023

AutoGraph: Optimizing DNN Computation Graph for Parallel GPU Kernel Execution

AAAI 2023technical

Deep learning frameworks optimize the computation graphs and intra-operator computations to boost the inference performance on GPUs, while inter-operator parallelism is usually ignored. In this paper, a unified framework, AutoGraph, is proposed to obtain highly optimized computation graphs in favo…

Cited by 6SourcePDFScholar
2023

Few-shot Joint Multimodal Aspect-Sentiment Analysis Based on Generative Multimodal Prompt

ACL 2023findings

We have witnessed the rapid proliferation of multimodal data on numerous social media platforms. Conventional studies typically require massive labeled data to train models for Multimodal Aspect-Based Sentiment Analysis (MABSA). However, collecting and annotating fine-grained multimodal data for MAB…

2023

S3IM: Stochastic Structural SIMilarity and Its Unreasonable Effectiveness for Neural Fields

ICCV 2023poster

Recently, Neural Radiance Field (NeRF) has shown great success in rendering novel-view images of a given scene by learning an implicit representation with only posed RGB images. NeRF and relevant neural field methods (e.g., neural surface representation) typically optimize a point-wise loss and make…

Cited by 37PDFScholar
2023

Uncertainty Guided Label Denoising for Document-level Distant Relation Extraction

ACL 2023long

Document-level relation extraction (DocRE) aims to infer complex semantic relations among entities in a document. Distant supervision (DS) is able to generate massive auto-labeled data, which can improve DocRE performance. Recent works leverage pseudo labels generated by the pre-denoising model to r…

2022

Context-Based Contrastive Learning for Scene Text Recognition

AAAI 2022technical

Pursuing accurate and robust recognizers has been a long-lasting goal for scene text recognition (STR) researchers. Recently, attention-based methods have demonstrated their effectiveness and achieved impressive results on public benchmarks. The attention mechanism enables models to recognize scene…

Cited by 62SourcePDFScholar
2022

PCL: Proxy-Based Contrastive Learning for Domain Generalization

CVPR 2022poster

Domain generalization refers to the problem of training a model from a collection of different source domains that can directly generalize to the unseen target domains. A promising solution is contrastive learning, which attempts to learn domain-invariant representations by exploiting rich semantic…

Cited by 157PDFcodeScholar
2021

Fast and Efficient DNN Deployment via Deep Gaussian Transfer Learning

ICCV 2021poster

Deep neural networks (DNNs) have been widely used recently while their hardware deployment optimizations are very time-consuming and the historical deployment knowledge is not utilized efficiently. In this paper, to accelerate the optimization process and find better deployment configurations, we pr…

Cited by 7PDFScholar
2020

DiffTaichi: Differentiable Programming for Physical Simulation

ICLR 2020poster

We present DiffTaichi, a new differentiable programming language tailored for building high-performance differentiable physical simulators. Based on an imperative programming language, DiffTaichi generates gradients of simulation steps using source code transformations that preserve arithmetic inten…

Cited by 486SourceScholar
2020

Multi-scale Two-way Deep Neural Network for Stock Trend Prediction

IJCAI 2020poster

Stock Trend Prediction(STP) has drawn wide attention from various fields, especially Artificial Intelligence. Most previous studies are single-scale oriented which results in information loss from a multi-scale perspective. In fact, multi-scale behavior is vital for making intelligent investment dec…

2019

Learning to Reconstruct 3D Manhattan Wireframes From a Single Image

ICCV 2019oral

From a single view of an urban environment, we propose a method to effectively exploit the global structural regularities for obtaining a compact, accurate, and intuitive 3D wireframe representation. Our method trains a single convolutional neural network to simultaneously detect salient junctions a…

Cited by 83PDFcodeScholar