This is the canonical public Adja speech dataset for the May 2026 thesis
release. It is intended for automatic speech recognition, text-to-speech, and
speech pipeline experiments.
This component duplicates the Orpheus speech source from JosueG/adja-tts-orpheus.
The source repository is treated as read-only provenance; this dataset repo is
the public canonical release surface.