CLEAR Global's Gamayun kits are a starting point for developing audio and text corpora for languages without pre-existing data resources. We create parallel data for a language by translating a pre-compiled set of general-domain sentences in English. If audio data is needed, these translated sentences are recorded by native speakers.
To scale corpus production, we offer four dataset versions:
Mini-kit of 5,000 sentences (kit5k)