International audience
While machine-generated data from marine sensors come with consistent pre-defined metadata, the logging of marine biological or biogeochemical sample data often involves human processing, leading to free-text descriptions provided as (raw) metadata. We propose here perspectives to optimise data management and better meet FAIR principles by leveraging AI and robust marine metadata norms, aiding from data originators to data managers in efficient annotation. We present a framework for interactive data validation prior to sharing through dedicated repositories, aligning with other data collections, and simplifying the archiving process. This will be facilitated primarily through the application of Natural Language Processing (NLP), aimed at identifying metadata, especially for the wide range of conventional and less common biogeochemical data (substances, matrices, etc.). While the marine biological community already uses tools like GBIF's Integrated Publishing Toolkit (IPT) and the Darwin Core (and its various extensions) respectively. The framework presented in this abstract complements these practices by assisting data providers in submission processes, maintaining FAIRness, and improving interoperability with minimal resource costs terms of human and computational resources.