
Building detection and instance segmentation from aerial imagery play a fundamental role in urban planning, disaster management, infrastructure monitoring, and smart city development. Although deep learning-based methods have achieved remarkable performance in building footprint extraction, most existing studies focus primarily on binary building detection while overlooking construction stage analysis. Furthermore, models trained on datasets from developed regions often exhibit limited generalization capability when applied to developing environments characterized by diverse construction patterns and heterogeneous urban layouts. To address these limitations, this study proposes a deep learning-driven framework for automated building detection, instance segmentation, and construction stage classification using high-resolution aerial imagery. The proposed framework employs Mask R-CNN integrated with ResNet-50 and ResNet-101 backbone networks to perform simultaneous object localization and pixel-level segmentation while categorizing buildings into three construction stages: complete, incomplete, and foundation. Experiments were conducted on a high-resolution aerial imagery dataset collected from Zanzibar Island, Tanzania, containing both urban and rural settlement structures. Experimental results demonstrate that the ResNet-101 backbone achieves superior performance, obtaining a building detection accuracy of 98% and a segmentation mean Average Precision (mAP) of 84.7%, compared with 95% detection accuracy and 76.4% mAP achieved by ResNet-50. The findings indicate that deeper residual architectures provide enhanced multi-scale feature representation and improved segmentation capability in complex aerial environments. The proposed framework offers an efficient and scalable solution for automated visual analysis of residential structures and has promising applications in urban monitoring, disaster assessment, infrastructure planning, and virtual geographic environment analysis.