
Concrete400 is a multi-task image dataset developed for automated visual inspection of concrete infrastructure. The dataset includes 379 full-size crack images with manually prepared pixel-level binary crack masks and crack-region bounding-box annotations prepared using Adobe Photoshop. It also includes 1,502 image patches for binary crack classification, comprising 633 crack patches and 869 no-crack patches.
The release is organised into three task-oriented components:
The segmentation images, masks, and object-detection coordinates are standardised to a common resolution of 512 × 512 pixels. The recommended reference partition contains 303 training pairs (image_001–image_303) and 76 validation pairs (image_304–image_379). Under this rule, crack pixels account for 3.07% of image pixels on average, with a median of 2.51% and a range of 0.40%–15.01%.
The photographs preserve realistic variations in illumination, shadows, concrete texture, surface condition, background clutter, and acquisition time. These field conditions, together with the sparse and irregular geometry of cracks, make the dataset suitable for evaluating robust computer-vision methods under realistic class imbalance.
Concrete400 can be reused for crack classification, semantic segmentation, instance segmentation, object detection, weakly supervised and semi-supervised learning, self-supervised representation learning, and crack-image generation. The dataset is intended for research in structural health monitoring, infrastructure inspection, civil engineering, and computer vision.