CATALYST: Cognitive-To-Autonomy-Inspired Two-Stage Training Data Generation with Local-System-Aware Selection Technique
Abstract
In conventional learning-based robotic dynamics modeling, physical information is mostly incorporated into the model or loss function, while the design of training data often relies on random sampling or uniform coverage, which can limit performance. To address this gap, this paper proposes the CATALYST framework, which generates optimal training data based on physics priors and the modeling structure of the chosen learning model. Stage 1 uses the CAD-derived inertia matrix M(q) to approximate the joint distribution of [q, M] with a PLM, thereby identifying the optimal locations for the local model centers (mu_k^{{opt}}). Stage 2 then optimizes an Operating-Point-Centered Excitation Trajectory (OPCET). This optimization simultaneously (i) aligns the trajectory with the target operating points (l_m), (ii) enforces range-of-motion (RoM) constraints (l_r), and (iii) achieves desirable velocity–acceleration statistics (large volume, isotropy, low correlation, captured by l_s). We validate the approach in simulation using a 3-DoF yaw–pitch–pitch manipulator, which allows visual demonstration of the process and outcomes. We then analyze the framework step by step. Results show that each stage meets its objective. A PLM trained on data generated by the proposed trajectories outperforms baselines (Spread/RoM, ill‑centered, Tukey‑windowed chirp, and cubic) in both torque regression and control. Thus, CATALYST yields more accurate regression and more reliable feedforward control than conventional designs.