PHONOSEMANTIC LAYER IN LARGE LANGUAGE MODELS

Authors

  • Mahmudjon Kuchkarov Author

Keywords:

Keywords: large language models; phonosemantics; sound symbolism; Odam Tili; SIGN–PHONE–SENSE; embodied cognition; biological grounding; language evolution; OT-PSL; cross-linguistic universals

Abstract

Large language models (LLMs) — including GPT-4, Claude, and Gemini — represent the most powerful artificial linguistic systems yet constructed, yet they share a fundamental architectural assumption: that the relationship between phonological form and meaning is entirely arbitrary, learned exclusively through distributional co-occurrence in text. This article argues that this assumption is empirically false and architecturally consequential. Drawing on the Odam Tili (Human Language) theory developed by Dr. Mahmudjon Kuchkarov, and on a longitudinal multi-site experimental dataset (N = 26; Fergana, Uzbekistan, 2003; New York, USA, 2011; Orlando, Florida, USA, 2017), we present systematic evidence for a biologically grounded phonosemantic layer in human language that current LLMs cannot acquire from text. Five phoneme–stimulus mappings are demonstrated: /h/ indexes physical load and hand agency (76% consistency across sites); /a/ indexes biological origin and distress (universal birth-cry datum; first letter in seven major writing systems); /o/–/u/ indexes long-distance projection (18% predominance in competitive shouting trials); /t/ indexes large-object impact (81%; preserved as a cross-family tree root: English tree, Uzbek tol, Russian topol, Hebrew tapuah); and /s/–/ʃ/ indexes smooth movement (63%/37% in snake-association surveys; parallel English–Uzbek /s/-initial vocabulary). We propose the Odam Tili Phonosemantic Layer (OT-PSL) as a concrete, biologically grounded architectural addition to current LLM design and discuss implications for cognitive science, language evolution, and neurolinguistics.

References

Anthropic. (2024). Claude 3 model card and system prompt. Anthropic Technical Report. https://www.anthropic.com/research

Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT '21), 610–623. https://doi.org/10.1145/3442188.3445922

Blasi, D. E., Wichmann, S., Hammarström, H., Stadler, P. F., & Christiansen, M. H. (2016). Sound-meaning association biases evidenced across thousands of languages. Proceedings of the National Academy of Sciences, 113(39), 10818–10823. https://doi.org/10.1073/pnas.1605782113

Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., & Amodei, D. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877–1901.

Corballis, M. C. (2002). From hand to mouth: The origins of language. Princeton University Press.

Cwiek, A., Fuchs, S., Draxler, C., Asu, E. L., Dediu, D., Hiovain, K., Kawahara, S., Koutalidis, S., Krifka, M., Lippus, P., Lupyan, G., Oh, G. E., Paul, J., Petrone, C., Ridouane, R., Reiter, S., Schümchen, N., Szalontai, Á., Ünal-Logacev, Ö., Zygis, M., & Winter, B. (2022). The bouba/kiki effect is robust across cultures and writing systems. Philosophical Transactions of the Royal Society B: Biological Sciences, 377(1841), Article 20200390. https://doi.org/10.1098/rstb.2020.0390

de Saussure, F. (1916). Cours de linguistique générale [Course in general linguistics]. Payot.

de Varda, A. G., & Strapparava, C. (2022). Phonosemantic correspondences are pretrainable. Proceedings of the 29th International Conference on Computational Linguistics (COLING 2022), 4213–4218.

Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL-HLT), 4171–4186. https://doi.org/10.18653/v1/N19-1423

Dingemanse, M., Blasi, D. E., Lupyan, G., Christiansen, M. H., & Monaghan, P. (2015). Arbitrariness, iconicity, and systematicity in language. Trends in Cognitive Sciences, 19(10), 603–615. https://doi.org/10.1016/j.tics.2015.07.013

Firth, J. R. (1957). A synopsis of linguistic theory 1930–55. In F. R. Palmer (Ed.), Selected papers of J. R. Firth 1952–59 (pp. 168–205). Longman.

Harnad, S. (1990). The symbol grounding problem. Physica D: Nonlinear Phenomena, 42(1–3), 335–346. https://doi.org/10.1016/0167-2789(90)90087-6

Harris, Z. S. (1954). Distributional structure. Word, 10(2–3), 146–162. https://doi.org/10.1080/00437956.1954.11659520

Hauser, M. D., Chomsky, N., & Fitch, W. T. (2002). The faculty of language: What is it, who has it, and how did it evolve? Science, 298(5598), 1569–1579. https://doi.org/10.1126/science.298.5598.1569

Hinton, L., Nichols, J., & Ohala, J. J. (Eds.). (1994). Sound symbolism. Cambridge University Press.

Imai, M., & Kita, S. (2014). The sound symbolism bootstrapping hypothesis for language acquisition and language evolution. Philosophical Transactions of the Royal Society B: Biological Sciences, 369(1651), Article 20130298. https://doi.org/10.1098/rstb.2013.0298

Kantartzis, K., Imai, M., & Kita, S. (2011). Japanese sound-symbolism facilitates word learning in English-speaking children. Cognitive Science, 35(3), 575–586. https://doi.org/10.1111/j.1551-6709.2010.01169.x

Kiefer, M., & Pulvermüller, F. (2012). Conceptual representations in mind and brain: Theoretical developments, current evidence and future directions. Cortex, 48(7), 805–825. https://doi.org/10.1016/j.cortex.2011.04.006

Kilpatrick, A. (2023). What artificial intelligence might teach us about the origin of human language. arXiv preprint arXiv:2301.06211. https://doi.org/10.48550/arXiv.2301.06211

Köhler, W. (1929). Gestalt psychology. Liveright.

Kuchkarov, M., & Kuchkarov, M. (2025). Human language as natural coding: The natural genesis of human language. World Scientific Research Journal, 36(1), 143–145.

Kuchkarov, M., Kuchkarov, M., & Sobirjonova, M. (2026a). Archaeology of language: Phonosemantic foundations and the universal water code. Journal of Applied Science and Social Science, 16(03), 417–421. https://www.internationaljournal.co.in/index.php/jasass/article/view/3705

Kuchkarov, M., Kuchkarov, M., & Sobirjonova, M. (2026b). The archaeology of natural coding: A phonosemantic deconstruction of the English lexicon through Odam Tili theory. International Multidisciplinary Journal for Research & Development, 13(03), 398–405. https://www.ijmrd.in/index.php/imjrd/article/view/5320

Lake, B. M., Ullman, T. D., Tenenbaum, J. B., & Gershman, S. J. (2017). Building machines that learn and think like people. Behavioral and Brain Sciences, 40, Article e253. https://doi.org/10.1017/S0140525X16001837

Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013). Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781. https://doi.org/10.48550/arXiv.1301.3781

Perniss, P., & Vigliocco, G. (2014). The bridge of iconicity: From a world of experience to the experience of language. Philosophical Transactions of the Royal Society B: Biological Sciences, 369(1651), Article 20130300. https://doi.org/10.1098/rstb.2013.0300

Pulvermüller, F., Hauk, O., Nikulin, V. V., & Ilmoniemi, R. J. (2005). Functional links between motor and language systems. European Journal of Neuroscience, 21(3), 793–797. https://doi.org/10.1111/j.1460-9568.2005.03900.x

Ramachandran, V. S., & Hubbard, E. M. (2001). Synaesthesia: A window into perception, thought and language. Journal of Consciousness Studies, 8(12), 3–34.

Rizzolatti, G., & Arbib, M. A. (1998). Language within our grasp. Trends in Neurosciences, 21(5), 188–194. https://doi.org/10.1016/S0166-2236(98)01260-0

Tomasello, M. (2008). Origins of human communication. MIT Press.

Wierzbicka, A. (1992). The semantics of interjection. Journal of Pragmatics, 18(2–3), 159–192. https://doi.org/10.1016/0378-2166(92)90050-L

Shtyrov, Y., Hauk, O., & Pulvermüller, F. (2004). Distributed neuronal networks for encoding category-specific semantic information: The mismatch negativity to action words. European Journal of Neuroscience, 19(4), 1083–1092. https://doi.org/10.1111/j.1460-9568.2004.03126.x

Published

2026-06-07

How to Cite

Mahmudjon Kuchkarov. (2026). PHONOSEMANTIC LAYER IN LARGE LANGUAGE MODELS. JOURNAL OF NEW CENTURY INNOVATIONS, 102(1), 332-350. https://journalss.org/index.php/new/article/view/32758