Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

MeSH2Wikidata: A set of tools for the interaction between MeSH keywords, OBO Foundry, and Wikidata for enriching biomedical knowledge

Domaine:

healthcare

Type de record:

softwaredataset
Créateur:
HouKhaBonChr
Hôte:avatar

The work consists of tools for the interaction between Wikidata and OBO Foundry and source codes for the use of MeSH keywords of PubMed publications for the enrichment of biomedical knowledge in Wikidata. This work is funded by the Adapting Wikidata to suppor… Project within the framework of the Wikimedia Foundation Resear….

To cite the work: Turki, H., Chebil, K., Dossou, B. F. P., Emezue, C. C., Owodunni, A. T., Hadj Taieb, M. A., & Ben Aouicha, M. (2024). A framework for integrating biomedical knowledge in Wikidata with open biomedical ontologies and MeSH keywords. Heliyon, 10(19), e38488. doi:10.1016/j.heliyon.2024.e384….

Wikidata-OBO

  • tool1.py: A tool for the verification of the semantic alignment between Wikidata and OBO ontologies.
  • frame.py: The layout of Tool 1.
  • tool2.py: A tool for extracting Wikidata relations between OBO ontology items.
  • frame2.py: The layout of Tool 2.
  • tool3.py: A tool for extracting multilingual language data for OBO ontology items from Wikidata.
  • frame4.py: The layout of Tool 3.

Wikidata-MeSH

  • correct_mesh2matrix_dataset.py: A source code for turning SisonkeBiotik-Africa/MeSH2M… into a smaller dataset for the biomedical relation classification based on the MeSH keywords of PubMed publications, named MiniMeSH2Matrix.
  • build_numpy_dataset.py: A source code for building the numpy files for MiniMeSH2Matrix (Relation type-based classification).
  • label_encoded.csv: A table for the conversion of Wikidata Property IDs into MeSH2Matrix Class IDs.
  • new_encoding.csv: A table for the conversion of Wikidata Property IDs into MiniMeSH2Matrix Class IDs.
  • super_classes_new_dataset_labels.npy: The NumPy File of the labels for the superclass-based classification.
  • new_dataset_labels.npy: The NumPy File of the labels for the relation type-based classification.
  • new_dataset_matrices.npy: The Numpy File of the MiniMeSH2Matrix matrices for biomedical relation classification.
  • first_level_new_data.json: The JSON File for the conversion of relation types to superclasses.
  • build_super_classes.py: A source code for building the numpy files for MiniMeSH2Matrix (Superclass-based classification).
  • FC_MeSH_Model_57_New_Data.ipynb: A Jupyter Notebook for training a Dense Model to perform the relation type-based classification.
  • FC_MeSH_Model_57_New_Data_SuperClasses.ipynb: A Jupyter Notebook for training a Dense Model to perform the superclass-based classification.
  • new_data_best_model_1: A stored edition of the best model for the relation type-based classification.
  • new_data_super_classes_best_model_1: A stored edition of the best model for the superclass-based classification.
  • MiniMeSH2Matrix_SuperClasses_Confusion_Matrix.ipynb: A Jupyter Notebook for generating the confusion matrix for the superclass-based supervised classification.
  • MiniMeSH2Matrix_Supervised_Classification_Agreement.ipynb: A Jupyter Notebook for generating the matrix of agreement between the accurate predictions for superclass-based classification and the ones for relation type-based classification.
  • Adding_References_to_Wikidata.ipynb: A Jupyter Notebook to identify the PubMed ID of relevant references to unsupported Wikidata statements between MeSH terms.
  • MeSH_Statistics.xlsx: Statistical data about MeSH-based items and relations in Wikidata.
  • ref_for_unsupported_statements.csv: Retrieved Relevant PubMed References for 1k unsupported Wikidata statements.
  • evaluate_pubmed_ref_assignment.ipynb: A Jupyter Notebook that generates statistics about reference assignment for a sample of 1k unsupported statements.
  • MeSH_Verification.xlsx: A list of inaccurate or duplicated MeSH IDs in Wikidata, as of August 8th, 2023.
  • WikiRelationsPMI.csv: A list of PMI values for the semantic relations between MeSH terms, as available in Wikidata.
  • WikiRelationsPMIDistribution.xlsx: Distribution of PMI values for all Wikidata relations and for specific Wikidata relation types.
  • WikiRelationsToVerify.xlsx: Wikidata relations needing attention because they involve Wikidata items with inaccurate MeSH IDs, they cannot be found in PubMed, or their PMI values are below the threshold of 2.
  • Mesh_part1.py: A Python code that verifies the accuracy of the MeSH IDs for the Wikidata items.
  • MeshWikiPart.py: A Python code that computes the pointwise mutual information values for Wikidata relations between MeSH keywords based on PubMed.
  • Demo.ipynb: A demo of the MeSH-based biomedical relation validation and classification in French.
  • Id_Term.json: A dict of Medical Subject Headings labels corresponding to MeSH Descriptor ID.
  • dict_mesh.json: Number of the occurrences of MeSH keywords in PubMed.
  • finalmatrix.xlsx: Matrix of PMI values between the 5k most common MeSH Keywords.
  • finalmatrixrev.pkl: Pickle File Edition of the PMI matrix.
  • pmi2.xlsx: List of significant PMI associations between the 5k most common MeSH Keywords reaching a threshold of 2.
  • Generate5kMatrix.py: A Python code that generates the PMI matrix.
  • clean_pmi2.py: A Python code to remove the relations already available in Wikidata from pmi.xlsx.
  • missing_rels.xlsx: The final list of the significant PMI associations that do not exist in Wikidata.
  • item_category.json: A dict for MeSH tree categories corresponding to MeSH items.
  • item_categorization.py: A Python code that generates a dict for MeSH tree categories corresponding to MeSH items.
  • classification.py: A Python code for classifying PMI-generated semantic relations between the most common MeSH Keywords.
  • results.xlsx: The output of the classification of the PMI-generated semantic relations between the most common MeSH Keywords.
  • ClassificationStats.ipynb: A Jupyter Notebook for generating statistical data about the classification.


Visit

figshare.com

Tasks

information extraction

Tags

Knowledge representation and reasoningData mining and knowledge discoveryInformation retrieval and web searchInformation modelling, management and ontologiesInformation systems organisation and managementKnowledge and information managementDigital curation and preservationDeep learningMachine LearningWikidata+14

Licenses

GPL 3.0+

Similaires

Wikidata as a Knowledge Base and Research Tool for Open Science in ArchaeologyESCoBox: A Set of Tools for Mini-Grid Sustainability in the Developing WorldA Metropolis Approach for Mesh Router Nodes placement in Rural Wireless Mesh NetworksKGARevion: An AI Agent for Knowledge-Intensive Biomedical QAMining Wikidata for Name Resources for African LanguagesMining Wikidata for Name Resources for African Languages

Wikidata as a Knowledge Base and Research Tool for Open Science in Archaeology

Abstract for the talk given at the Digital Archaeology Bern 2023 conference,

ESCoBox: A Set of Tools for Mini-Grid Sustainability in the Developing World

Mini-grids powered by photovoltaic generators or other renewable energy sources have the potential t

A Metropolis Approach for Mesh Router Nodes placement in Rural Wireless Mesh Networks

Wireless mesh networks appear as an appealing solution to reduce the digital divide between rural an

KGARevion: An AI Agent for Knowledge-Intensive Biomedical QA

Biomedical reasoning integrates structured, codified knowledge with tacit, experience-driven insight

Mining Wikidata for Name Resources for African Languages

This work supports further development of language technology for the languages of Africa by providing a Wikidata-derived resource of name lists corresponding to common entity types (person, location, and organization). While we are not the first to mine Wikidata f

Mining Wikidata for Name Resources for African Languages

This work supports further development of language technology for the languages of Africa by providing a Wikidata-derived resource of name lists corresponding to common entity types (person, location, and organization). While we are not the first to mine Wikidata f