Introduction: Music and language are both found in all societies. But did humans evolve to make music, or is our musicality a byproduct of language evolution? Surprisingly, there remains little data on which aspects of music and language are shared or distinct across cultures, due to the limitations of automated acoustic analysis and the lack of cross-cultural datasets with human annotations. Approach: We combined objective acoustic analysis with subjective speaker annotations to compare recordings of singing, lyrics recitation, spoken description of the song, and the instrumental version across diverse languages. Findings: Pilot analyses of recordings in five languages (Japanese, English, Yoruba, Farsi, and Marathi) suggest that tempo is the feature that most strongly differentiates speech and song (song is slower), while melodic interval is most similar (both speech and song are characterized by small intervals <700 cents). Discussion: We will use a Registered Report framework to rigorously test candidate features on a sample of over 50 languages, revealing cross-culturally shared and distinct features of speech and song, and shedding light on the evolution of language and music.