Abstract
Corporate budget-revenue allocation and labor cost expenditure planning are tightly coupled decisions, yet prevailing practice treats them in isolation. This paper proposes HMARL-BW, an uncertainty-aware hierarchical multi-agent deep reinforcement learning framework that unifies budget allocation and workforce cost planning within a constrained Markov decision process. An upper-level agent apportions the cross-departmental budget pool, while lower-level agents optimize headcount and compensation under budgetary constraints; inter-departmental dependencies are encoded via a graph neural network. Distributionally robust risk constraints and a scenario simulator based on a conditional variational autoencoder (CVAE) enhance robustness against demand shocks, and attribution analysis based on Shapley additive explanations (SHAP) improves decision transparency. Experiments on a CSMAR-calibrated simulation environment and a digital-twin simulator show that the proposed method outperforms linear programming (LP), genetic algorithm (GA), MAPPO, and single-agent reinforcement learning (SA-RL) baselines, improving the revenue growth rate and labor cost efficiency by 11.8% and 9.4%, respectively, and reducing budget execution deviation by 23.3%, relative to the strongest baseline (SA-RL), while retaining 80.2% of the nominal reward under a -40% demand shock versus 64.1% for SA-RL.