Agricultural scene understanding through semantic segmentation is fundamental for precision agriculture, enabling phenological analysis, yield estimation, disease assessment, and crop management. Despite recent advances in CNN-, transformer-, and Mamba-based architectures, existing methods still exhibits inherent limitations in balancing fine-grained local details and global semantic reasoning. We propose DeMamba, a novel architecture that integrates deformable convolution-based local feature modeling with efficient Mamba-based global context modeling. Specifically, we introduce a Global Semantic State Validation (GSSV) module to enhance long-range dependencies and an Adaptive Local–Global Fusion (ALGF) gated mechanism for dynamic feature integration. Furthermore, a two-stage training strategy with gate supervision and structure-aware stage loss is employed to improve segmentation consistency. Extensive experiments demonstrate that DeMamba consistently outperforms existing SOTA methods on a challenging UAV-based rice panicle dataset (+1.00 IoU, +2.13 Acc, and +5.04 to +3.36 IoU across tiny-to-large panicles), while also improving performance on additional agricultural benchmarks, confirming its robustness and generalization.