Musilanguage and the Song Beneath Our Words
The Sound Before the Word
Across the traditions explored on this site — the vibrating consonants of Egyptian temple names, the sung hexameters of Greek prophecy, the drum and voice of Indigenous ceremony — one thing keeps surfacing. Sacred speech is rarely just speech. It leans toward melody. Pitch carries meaning. Rhythm carries intention. Somewhere underneath the world's chant, mantra, and liturgy, sound and sense seem to share a single root.
It turns out this isn't only a spiritual intuition. It's also a live question in the science of human origins. A number of researchers who study the evolution of communication have proposed that music and language were not always two separate things — that both may have branched from one undivided system of expressive sound. Two names dominate this conversation: the musicologist Steven Brown, who calls that shared ancestor musilanguage,1 and the archaeologist Steven Mithen, who calls it Hmmmmm.2
One Root, Two Branches
Brown's musilanguage model treats music and language as "reciprocal specializations" of a single ancestral system — one that carried both referential meaning (what we'd now call language) and emotive meaning (what we'd now call music) in the same breath.1 For this precursor stage to function as a scaffold for both descendants, Brown argues it needed at least three ingredients: lexical tone, the use of pitch to help distinguish meaning; combinatorial phrase formation, the ability to string units together into longer utterances; and expressive phrasing, the shaping of a vocal line for emotional effect.1 In other words, before there were words, there was already tone, phrase, and feeling moving together.
Mithen's account, laid out at length in The Singing Neanderthals, arrives at a similar place from a different angle — through the fossil and archaeological record of early hominids rather than through comparative linguistics. He proposes that a communication system existed long before modern language, one carrying "the characteristics that are now shared by music and language, but that split into two systems at some date in our evolutionary past."3 He names this system with the intentionally sing-song acronym Hmmmmm: Holistic, Manipulative, Multi-modal, Musical, and Mimetic.3
Each term describes a different layer of that pre-linguistic communication. Holistic means the utterances weren't built from separable words the way a sentence is — a whole phrase carried a whole meaning, the way a sigh or a lullaby line does now. Manipulative means these calls were used to shift another person's emotional state or behavior, not simply to report facts. Multi-modal means voice was inseparable from gesture and movement — the body was part of the utterance. Musical means pitch, rhythm, tempo, and repetition were doing real communicative work. And mimetic means sound symbolism and imitation — echoing the shape of the thing being communicated — played a central role.3 Put together, Mithen's Hmmmmm looks less like proto-speech and more like proto-song: expressive, embodied, tonal, and whole.
Wray's Holistic Phrase
A related idea comes from the linguist Alison Wray, whose work on what she calls "holistic protolanguage" fed directly into Mithen's model. Where later theories of language evolution often imagine early humans building up meaning from small parts — sounds into words, words into grammar — Wray's holistic view runs the other direction. The earliest utterances, in this account, were unsegmented wholes; the breaking-apart into discrete words came later, as a kind of decomposition of something that started out unified.4 It's a strikingly different picture of where language comes from: not built up from bricks, but poured out whole and only gradually carved into pieces.
Why This Matters for Sacred Sound
None of this proves that ritual chant is a literal echo of Pleistocene vocal behavior — the evolutionary record for something as intangible as early communication is famously thin, and musilanguage remains a contested hypothesis rather than settled fact.1 But it offers something valuable to anyone drawn to sacred sound traditions: a serious, evidence-engaged account of why chant, mantra, and tonal prayer feel so native to us. If Brown and Mithen are right, we aren't layering music onto language when we chant a name or intone a syllable. We're returning, if only briefly, to something closer to where both came from — a mode of communication that was never purely referential and never purely musical, but both at once.
This is, in a way, what devotional traditions across the world seem to have rediscovered independently and repeatedly. The tonal precision required in reciting Quranic Arabic. The belief, running through ancient Egyptian ritual, that the correctly voiced name of a thing had power over the thing itself. The chant and drum of Indigenous ceremony, where the body moves with the voice rather than beside it. Each of these traditions, in its own vocabulary, treats sound as more than a container for information — treats it as holistic, embodied, and capable of shifting the state of the listener. That is Hmmmmm's own definition, translated into ritual.3
There's a quiet invitation in this research, whether or not one reads it as literal evolutionary history. It suggests that reaching for tone before words — a hum, a drone, a single sustained syllable — isn't a retreat from meaning. It may be a return to the oldest form meaning ever took.
What This Feels Like in Practice
Spend time in a sound healing session, or simply sit with a drone or a bowl for a few minutes, and something in Brown and Mithen's model becomes easy to recognize. A single sustained tone doesn't need translating. It doesn't ask to be parsed the way a sentence does. It arrives holistically, the way Mithen's model says the earliest utterances did — whole, felt, and immediate, rather than assembled piece by piece.3 That's part of why tone-based practices can reach a person who is agitated, grieving, or otherwise unable to take in words. The nervous system doesn't need language to recognize a falling pitch as a settling one, or a steady pulse as an invitation to slow down. It already knows. That recognition may be older than any word we have for it.
This is also, in a smaller way, why so many contemplative traditions treat the correct sounding of a syllable as inseparable from its meaning — why a mantra is chanted rather than read, why a name of God is intoned rather than merely spoken, why prayer so often becomes song under pressure. If Brown and Mithen are right, that instinct isn't decorative. It's a return route to a mode of communication that never separated sense from sound in the first place, one where the message and the music were, for a very long stretch of our history, the same thing.
None of this requires believing every detail of the musilanguage hypothesis to be settled. The value, for those of us drawn to sacred sound, is simpler: it gives serious, well-argued language for something ritual practice has always assumed — that voice carries more than information, and that the oldest layer of human communication may still be the one doing the real work when words run out.
Sources
- ↩ Brown, S. (2001). "The 'Musilanguage' Model of Music Evolution." In N. L. Wallin, B. Merker, & S. Brown (Eds.), The Origins of Music (pp. 271–300). MIT Press. neuroarts.org/pdf/musilanguage.pdf
- ↩ Botha, R. (2009). "On musilanguage/'Hmmmmm' as an evolutionary precursor to language." Lingua. sciencedirect.com/science/article/abs/pii/S0271530908000025
- ↩ Mithen, S. (2005). The Singing Neanderthals: The Origins of Music, Language, Mind and Body. Weidenfeld & Nicolson. Summary reference: eamusic.dartmouth.edu — response essay on Mithen's Hmmmmm model
- ↩ Wray, A., cited in Botha, R. (2009), on "holistic protolanguage." sciencedirect.com/science/article/abs/pii/S0271530908000025