Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Integrating Cyc and Wikipedia: folksonomy meets rigorously defined common-sense

Record type:

datasetsoftware
Creator:
O MCat
Host:avatar
Integration of ontologies begins with establishing mappings between their concept entries. We map categories from the largest manually-built ontology, Cyc, onto Wikipedia articles describing corresponding concepts. Our method draws both on Wikipedia’s rich but chaotic hyperlink structure and Cyc’s carefully defined taxonomic and common-sense knowledge. On 9,333 manual alignments by one person, we achieve an F-measure of 90%; on 100 alignments by six human subjects the average agreement of the method with the subject is close to their agreement with each other. We cover 62.8% of Cyc categories relating to common-sense knowledge and discuss what further information might be added to Cyc given this substantial new alignment.

Visit

figshare.com

Tags

890399 Information Services not elsewhere classified080608 Information Systems Development Methodologies080105 Expert SystemsSchool of Humanities and Social Sciences

Licenses

All Rights Reserved

Similar

Common Sense Beyond English: Evaluating and Improving Multilingual Language Models for Commonsense ReasoningWikipedia: wikipedia-af (Afrikaans)Somaliska Wikipedia Somali WikipediaWikipediaWikipediaWikipedia

Common Sense Beyond English: Evaluating and Improving Multilingual Language Models for Commonsense Reasoning

Commonsense reasoning research has so far been limited to English. We aim to evaluate and improve popular multilingual language models (ML-LMs) to help advance commonsense reasoning (CSR) beyond English. We collect the Mickey Corpus, consisting of 561k sentences in

Wikipedia: wikipedia-af (Afrikaans)

Wikipedia is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia

Somaliska Wikipedia Somali Wikipedia

Korpus av somaliska Wikipedia Corpus of Somali Wikipedia

Wikipedia

Wikipedia dataset containing cleaned articles of all languages. The datasets are built from the Wikipedia dump (https://dumps.wikimedia.org/) with one split per language. Each example contains the content of one full Wikipedia article with cleaning to strip markdow

Wikipedia

Wikipedia dataset containing cleaned articles of all languages. The datasets are built from the Wiki

Wikipedia

Wikipedia dataset containing cleaned articles of all languages. The datasets are built from the Wiki