Uppsats

Synthetic Data-driven Inventory Control Optimisation via Reinforcement Learning

Kandidat-uppsats

Högskolan i Halmstad/Akademin för informationsteknologi

Publicerad: 2026

Språk: Engelska

Sammanfattning

This work investigates whether reinforcement learning can match or outperform a classical Standard Inventory Policy (SIP) in spare parts inventory management for the automotive aftermarket. Four reinforcement learning algorithms are evaluated against the single-echelon (s,Q) baseline using a synthetic simulator that replicates the cost calculation and ordering structure of an automotive aftermarket inventory system. The four algorithms compared are Proximal Policy Optimization, Soft Actor- Critic, Advantage Actor-Critic, and Twin Delayed DDPG, all instantiated through the Stable-Baselines3 library. A SIP-triggered continuous action mask is introduced to bound the agent’s order quantity within the feasible window. On a 439-day test period, Soft Actor-Critic with the mask reduces total cost by 2.5% relative to the SIP baseline while matching the same service level. Twin Delayed DDPG with the mask reaches the 0.95 service target, and the on-policy algorithms PPO and A2C fail to reach the baseline. The results show that reinforcement learning can outperform the classical inventory rule, but only conditional on both the algorithm choice and the application of a specific mask.

Information

Lärosäte / institution
Högskolan i Halmstad/Akademin för informationsteknologi
Publiceringsdatum
2026
Uppsatstyp
Kandidat-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.