IOSTOM: Offline Imitation Learning from Observations via State Transition Occupancy Matching
Offline Learning from Observations (LfO) focuses on enabling agents to imitate expert behavior using datasets that contain only expert state trajectories and separate transition data with suboptimal actions. This setting is both practical and critical in real-world scenarios where direct environment…