The sustainable intensification of agriculture requires precise water management strategies, particularly the adoption of deficit irrigation to cope with increasing global water scarcity. Optimizing the intraseasonal scheduling of limited water supplies represents a complex sequential decision-making problem governed by stochastic climatic boundary conditions and nonlinear crop responses. While deep reinforcement learning has emerged as a powerful paradigm for developing flexible closed-loop control policies, current applications often exhibit instability due to high-dimensional state spaces and the inadequate handling of physical constraints. This research advances the field by proposing and benchmarking a tailored deep reinforcement learning framework based on the Proximal Policy Optimization algorithm, coupled with the AquaCrop-OSPy soil-crop-atmosphere simulation model. The study introduces specific architectural enhancements to the learning agent, including a reduced observation space consisting of five causal biophysical variables, action masking to ensure strict adherence to seasonal water quotas, and a dense reward function based on transpiration efficiency to guide the learning process. To rigorously assess the algorithmic performance and quantify the value of information, the proposed approach is benchmarked against a comprehensive set of baselines, including a standard deep reinforcement learning implementation and a global evolutionary algorithm. Crucially, the evolutionary algorithm is configured with perfect foresight of future weather events, thereby serving as a theoretical oracle that defines the upper bound of achievable crop water productivity. Experimental validation was conducted for maize cultivation under deterministic conditions in Tunis and stochastic climate scenarios in Nebraska. The empirical results demonstrate that the proposed agent effectively navigates the trade-off between water conservation and yield maximization, capturing approximately 93.5 percent of the theoretical yield potential defined by the evolutionary algorithm oracle. This result indicates a minimal performance penalty attributed to the lack of future weather knowledge. In contrast, the reference implementation failed to converge to efficient policies under tight resource constraints. Furthermore, the economic evaluation highlights that the proposed strategy not only stabilizes yields during extreme drought years but also increases mean net profits by up to 66% compared to the reference baseline. These findings confirm that integrating domain knowledge through action masking and feature selection enables deep reinforcement learning to serve as a robust riskminimization tool, yielding near-optimal irrigation scheduling without requiring extensive weather forecasting.