
Inside a soundproof cabin in Tokyo, a juvenile male Bengalese finch settles on his perch, adjusts his throat, and begins to practice. It is day 65 of his life.
To the human ear, his song is a fluttering, slightly chaotic cascade of chirps, buzzes, and clicks, a tentative rehearsal of the elaborate display he will need in adulthood. On a computer screen outside the isolation chamber, the acoustic output scrolls past as a live sonogram: a space densely populated with vertical streaks and frequency glides, broken by tiny slivers of silence.
To an untrained ears, the bird’s performance sounds continuous and unformatted, a stream of acoustic fluid pouring into the air. But a closer look at the sonogram reveals discrete building blocks emerging.
Researchers call them syllables. A single Bengalese finch possesses a repertoire of about eight distinct syllable types, which it stitches together into extended vocal performances known as song bouts.
For decades, bio-acousticians assumed that learning a song was simply a matter of a chick listening to an adult tutor and memorising a rigid, linear sequence: A leads to B, B leads to C, C leads to D. But watching the juvenile finch over several weeks, a stranger reality started revealing itself to the researchers.

The chick does not copy his tutor’s song as an unbroken tape recording. Instead, he fractures the tutor’s song, sampling fragments, recombining local transitions, and reassembling the pieces into a custom repertoire.
He is not merely mimicking; he is parsing. This subtle act of parsing lies at the heart of one of the most remarkable discoveries in contemporary cognitive science.
The Experiment
In a landmark 2026 paper published in Science Advances, an international team of researchers, Simon Kirby, Kazuo Okanoya, Ellen C. Garland, Miki Takahasi, and Inbal Arnon, demonstrated that the songs of Bengalese finches contain a complex, language-like statistical architecture.
By applying an analytical pipeline originally designed to study how human infants discover words in spoken speech, the researchers revealed that birdsong is organised into statistically coherent ‘chunks’ of syllables whose frequency of occurrence follows a striking mathematical power law known as Zipf’s law.

This is an empirical mathematical pattern where the frequency of an item is inversely proportional to its rank in a list. In every human language, if you tally every word in a large body of text and line them up from most common to least common, a remarkably steady pattern appears.
The most common word shows up roughly twice as often as the second-most common, three times as often as the third, four times as often as the fourth, and so on. When these counts are plotted so that both the ranking and the frequency stretch out on an exponential scale, the points fall along a straight, steeply falling line. Linguists call this a Zipfian distribution, after George Kingsley Zipf, who brought it to wide attention in the mid-twentieth century.
Coming back to the 2026 Science paper, what makes this discovery startling is not merely that birdsong possesses structure, but that this precise statistical structure is shared by three evolutionary lineages separated by hundreds of millions of years: human speech, humpback whale song and songbird vocalisations.
None of these species inherited this structure from a common ancestor. Instead, all three appear to have arrived at the same organisational design through convergent evolution. Yet, as the researchers were quick to emphasise, this discovery raises a profound conceptual question. The mathematical tools used to dissect the finch’s song can identify statistical boundary lines in a stream of sound, but they cannot tell us how the bird itself experiences those sounds. Does the finch perceive its song as a chain of individual statistical links, or does it apprehend each multi-syllable phrase as an indivisible, holistic unit?
To answer that question, we must look not only to modern computational bio-acoustics, but also to a long-overlooked epistemological insight from the ancient world, one that redefines what it means to parse a sequence in time. To understand how Kirby and his colleagues revealed the hidden architecture of birdsong, begin with a fundamental puzzle of human language acquisition.
Adults do not insert tidy gaps between spoken words. Speech arrives as an unbroken acoustic stream: the textbook example is ‘Look at the pretty baby’ becomes ‘look at the pretty baby’. Long before infants grasp word meanings, they somehow parse this continuum into discrete units. How? In the mid-1990s, Jenny Saffran and colleagues showed that eight-month-old infants solve the problem through statistical learning. They track transitional probabilities, the likelihood that one syllable will follow another, and use those probabilities to detect word boundaries. In other words, the human infant’s brain acts as a silent statistical engine, continuously tracking predictability. Wherever predictability plunges, the brain posits a boundary line.
In their 2026 study, Kirby, Okanoya, Garland, Takahasi, and Arnon saw that the same method babies use to find words could open up animal communication without human bias. Rather than guessing where a bird’s song ‘words’ should begin and end, they let the bird’s own patterns of sound predictability draw the map.
They gathered song recordings from six young male Bengalese finches across several months of growth. Each bird was recorded at five key stages, starting around day 60, when the early songs became clear enough to study and continuing to day 120, when the song locked into its final adult form. Every chick had been raised with a single adult male tutor, giving the team a full picture of how the young birds absorbed their vocal heritage over time.
To parse the songs, the researchers measured how likely each sound was to follow the one before it. They avoided any fixed cut-off. Instead they looked for sudden drops in predictability: whenever the chance of the next sound fell to less than half the chance of the sound just before it, a boundary was marked.
The logic was simple. Inside a natural phrase, one sound reliably leads to the next. At the edge between phrases, that reliability collapses. When the method ran through the finch recordings, it sliced the continuous song streams into repeating groups of several notes. On average it made about eleven cuts in each bout, recovering a set of recurring multi-note phrases. The unbroken flow of birdsong had been turned into a clear collection of coherent units.
Once the continuous song had been broken into a catalogue of repeating phrases, the next question was straightforward: How often does each phrase appear? True Zipfian patterns are extremely uncommon outside human speech. Many animals use a few calls a lot and most calls rarely, yet almost none of their multi-sound sequences form a clean, reliable power-law curve of this kind.
When Kirby and his colleagues ranked the finch phrases by how often each one occurred, the picture was unmistakable. The counts followed a power-law curve with striking precision. The steepness of that curve, the power-law exponent, averaged 1.05 across the recordings. In quantitative linguistics, a value of 1.05 is virtually the same as the slope found in human languages ranging from English and Mandarin to Swahili and Basque.
The researchers found still another classic pattern of human language tucked inside the data: the shorter a phrase, the more often it appears. In everyday speech, common words are brief: ‘the’, ‘and’, ‘it’; uncommon ones stretch out— ‘antidisestablishmentarianism’, ‘extraterrestrial’.
The finch songs showed the same efficiency: the brief multi-note phrases turned up far more often than the longer ones. The birds were following a simple, widespread principle of communication: keep the most-used signals short so that effort is saved without losing the ability to convey a rich set of distinct messages. Common sense suggests that a young bird should start with simple, unorganised acoustic noise and gradually construct higher-order statistical structure over time. Under this bottom-up view, juvenile song should initially look statistically chaotic, showing a poor fit to a power law and only settle into a Zipfian distribution late in development as the song approaches its crystallized adult target.
The data told a completely different counter-intuitive story. When the team analysed the song recordings across developmental time, they discovered that the goodness of fit to a power law was essentially invariant. At day 60, when the young bird’s song is still slushy, unrefined, and phonetically distant from its tutor’s target, the rank-frequency distribution of its statistical chunks already fit a power law just as tightly as adult song.
The Evolutionary Convergence
The implications of the Bengalese finch study extends far beyond avian bio-acoustics. By demonstrating that finch song possesses statistically coherent chunks organized according to a Zipfian power law, Kirby and his team completed an extraordinary evolutionary triad.
In 2025, a parallel study led by Inbal Arnon and Simon Kirby applied the same infant-inspired segmentation pipeline to eight years of culturally transmitted humpback whale songs recorded across the Pacific Ocean. The whale songs yielded the exact same statistical profile: recurring multi-element motifs defined by transitional probability drops, organized in a Zipfian power-law distribution. The breadth of this comparative convergences in the vast evolutionary landscape is amazing to say the least.
Here are three lineages separated by hundreds of millions of years of independent evolutionary history. One is a terrestrial primate using vocal-tract articulation to convey abstract propositional thought; the second is a pelagic marine mammal generating low-frequency acoustic displays that sweep across ocean basins; the third is a small arboreal passerine bird using a specialized vocal organ (the syrinx) to perform rapid courtship displays.
What unites these three wildly disparate biological systems?
All three systems, human language, humpback whale song, and Bengalese finch song, rely on complex sequences that are not fixed by genes. They must be learned from others and handed down across generations. Children absorb language from their community. Male whales pick up shifting song patterns from neighbouring groups across the oceans. Young finches copy elaborate note combinations from adult tutors. In other words, this is cultural transmission.

Whenever an intricate sequence must pass through the limited mind of a learner, generation after generation, it meets a harsh filter. Patterns that are hard to break into pieces or hard to remember get scrambled or lost. Patterns that fit the brain’s natural tendency to find statistical chunks survive, stick, and get passed on cleanly. Over the evolutionary time this filter shapes the signals.
Complex sequential communication, whatever the species, tends to settle into the same basic design: coherent chunks whose frequencies follow a steep power-law curve. A handful of very common ‘anchor’ chunks act as reliable landmarks, making it far easier for a novice to carve the continuous stream of sound into recognizable pieces.
Caveat
Here sensationalism should be met with strict scientific caveat. Finding language-like statistics in birdsong does not mean finches possess language. Their chunks are not words with meanings as yet we know of. There is no evidence of symbols, concepts, or sentences that make statements. The researchers themselves stressed this point: the song serves mainly as a male display to attract mates. The chunks carry no dictionary definitions, no stories, and no declarative content. What the work reveals is more fundamental.
Statistical features long treated as unique signatures of human language are in fact general properties of any complex signal that must be learned and culturally transmitted. Structure can arise without meaning. Organisation can assemble itself without intention. The statistical method developed by Kirby and his colleagues is a remarkable tool. By detecting sudden drops in the predictability of one sound following another, it can locate hidden seams in a continuous acoustic stream.

The Sequence and the Whole
Physical sound exists only in time, one moment after another. Yet the cognitive unit it produces seems to exist all at once.
Consider a three-note bird phrase: A-B-C. The notes are produced in sequence. When C is sounding in the air, A has already vanished; it survives only as a fading memory. At no physical instant do A, B, and C exist together in the world. How, then, does the receiver’s mind gather these fleeting, non-overlapping events into one coherent, functional whole?
Modern cognitive science and bioacoustics still wrestle with this sequence-to-whole problem. Fifteen centuries earlier, a classical Indian philosopher framed an unusually precise answer to the same epistemological puzzle.
His name was Bhartṛhari (roughly 5th century CE), a grammarian and philosopher of language whose major work is the Vākyapadīya (‘Treatise on Sentences and Words’). He was not an experimental scientist and knew nothing of probability algorithms, neural networks, or birdsong. Yet he was intensely occupied with the exact tension that faces researchers today: the conflict between the temporal succession of physical sound and the holistic character of cognitive grasp.
Bhartṛharian Explanation
When Kirby and his colleagues run their segmentation pipeline across a corpus of birdsong, they are measuring dhvani, the physical, moment-by-moment stream of sound itself. The relative drops in transitional probability simply mark the places where predictability collapses inside that temporal sequence. This is a major empirical advance. It reveals the material conditions that make a complex signal learnable.

Yet Bhartṛhari’s framework issues a sharp warning against a classic category error: we must not mistake the statistical boundary drawn by an algorithm for the internal representation formed inside the bird’s brain.

The pipeline performs apoddhāra, an analytical extraction that slices the continuous song into discrete multi-syllable chunks on a computer screen. For the finch, however, those same sounds may not be experienced as additive strings of syllables merely glued together by high transitional probabilities. Instead, each successive syllable (dhvani) may leave a cumulative neural imprint (saṃskāra) in the avian auditory system.
These traces build until, at a critical moment, the complete phrase (sphoṭa) flashes into existence as a single, indivisible motor or perceptual whole – a gestalt grasped all at once rather than assembled piece by piece. This conceptual reframing helps explain one of the most puzzling findings of the Science Advances paper: the developmental trajectory of juvenile song.

Looking back at the experiment one will find that at day 60, when a young finch’s song is still rough and only loosely resembles its tutor’s, the frequency distribution of its emerging phrases already follows a clear Zipfian power law (R² ≈ 0.89). Under a purely bottom-up account this is surprising: the bird has not yet mastered clean transitions, yet the statistical signature of adult organisation is already visible. A Bhartṛharian perspective can explain this puzzle.
The juvenile brain may already contain a latent, holistic organisational principle, the structural attractor of the sphoṭa, from the outset of vocal practice. The bird does not construct sequential order from scratch out of isolated syllable fragments. Instead, over weeks of motor practice and social tutoring, it gradually clarifies the physical sound stream (dhvani). The tutor’s song serves as a guide that sharpens the internal transitions, turning a diffuse, uneven vocal display into a more precise and coherent expression of the target phrase. The modern information-theoretic view explains how cultural transmission, across generations, filters signals for learnability. The Bhartṛharian view offers an ontology for how a single organism integrates those sequences into unified perceptual or motor wholes.

The two accounts operate at different levels; they complement rather than compete.
Bhartṛharian Epistemology
Can empirical science then use this Bhartṛharian framework? Can bioacousticians move past corpus statistics on a screen and test whether animals actually represent statistical chunks as indivisible wholes in the brain? The path forward is to design experiments that force a direct contrast between local statistical tracking and global, holistic pattern recognition. Three concrete paradigms illustrate the approach.
- Global perturbation: Constructing synthetic sequences that keep every local transition probability intact, so that A still strongly predicts B, yet deliberately scramble the larger phrase architecture. If the animal processes song only as a chain of local probabilities, the altered sequences should register as normal. If the brain instead represents phrases as indivisible units, the global disruption should produce clear behavioural orientation responses or neural mismatch signals in auditory regions.
- Acoustic invariance and pattern completion: Train animals to recognise a specific target chunk, then present degraded versions, pitch-shifted, time-compressed, or with individual syllables masked by noise. If recognition depends on an abstract, holistic representation that remains stable despite changes in the physical surface of the sound, the animal should continue to identify the target even when local transitional cues are broken.
- Neural manifold dynamics: Using high-density probes in core avian auditory and vocal centres track the trajectory of neural population activity in real time while a bird hears song. If representation is strictly local and sequential, neural states should advance continuously and roughly linearly from syllable to syllable. If the brain constructs holistic gestalts, the population activity should show sharp, non-linear transitions – abrupt jumps in state space at critical moments when incomplete acoustic input triggers global pattern completion.
It does not matter whether Bhartṛhari is proved right or wrong. That is not how science works. But the framework is what is important. The very fact that Bhartṛharian framework has becomes a forceful stimulus for further exploratory approaches – a strong epistemological tool, is in itself shows how one can truly put to use IKS.
The work of Kirby, Okanoya, Garland, Takahasi, and Arnon marks a genuine advance.
By showing that Bengalese finches, humpback whales, and humans share statistically coherent chunks arranged in Zipfian distributions, it demonstrates that cultural transmission through social learning systematically shapes signals toward efficient, learnable forms. Its deeper value, however, lies in the conversation it opens. Computational bioacoustics supplies the empirical tools and quantitative precision needed to map the physical signal. Indic frameworks such as Bhartṛhari’s doctrine of sphoṭa supply the conceptual reminder that beneath every measured stream of sound is a living system trying to make coherent sense of time.
Journal Reference: Simon Kirby et al. , Language-like statistical structure arises in learned signaling: evidence from birdsong. Sci. Adv.12,eaea3015(2026). DOI:10.1126/sciadv.aea3015
AI disclosure: During the preparation of this work, the author has used Chatgpt, Gemini-Notebook and Grok for grammar corrections, polishing of the language and content editing. After using these tools, the author has reviewed and edited the content again as needed and takes full responsibility for the publication of the content.