Accurate estimation of sorghum (Sorghum bicolor [L.] Moench) grain yield remains a major bottleneck in breeding programs across sub-Saharan West Africa, where traditional methods rely on manual counting and weighing of grains. These approaches are labor-intensive, time-consuming, and prone to-induced variability, limiting throughput and reproducibility. This study presents an end-to-end high-throughput phenotyping pipeline for automated grain detection and mass estimation using computer vision. The workflow integrates smartphone-based image acquisition, automated foreground extraction, and grain detection using YOLOv11 models, followed by count-to-mass calibration. Three YOLOv11 architectures, small, medium, and large, were evaluated under identical training conditions using transfer learning from COCO pre-trained weights. Among the tested models, the medium configuration provided the best trade-off between precision and detection performance, with precision values reaching up to 0.861. A preprocessing step based on automatic background masking was applied prior to detection, significantly improving robustness under heterogeneous field conditions. Grain count was converted to mass using a linear calibration model. For two locally relevant sorghum elite lines, Faourou and Payenne, strong linear relationships were observed between detected grain number and measured grain weight (R² > 0.998), with mass coefficients k ≈ 0.026 g/grain. To facilitate adoption, a user-friendly R Shiny application was developed, allowing users to upload images, perform automated grain detection, and estimate grain mass using user-defined calibration coefficients. This pipeline provides a scalable and reproducible approach for rapid grain phenotyping and offers strong potential to accelerate selection decisions in sorghum breeding programs under field conditions.