The following is an excerpt from my masters thesis. This section attempts to demonstrate the key concerns of the thesis and motivation behind it. I have removed citations and references in this blog post, as I am not considering this specific post to be an actual academic work.

A Case Study in Lower Sorbian

Why is an exploration of capabilities of models in low resource languages important? Why might we want dedicated models for such languages? I will explain with the case of Lower Sorbian. Sorbian is a West Slavic language of the Lechitic family, which consists of Polish and a number of low speaker count and low resource languages. Sorbian is spoken in eastern Germany, and is split into the upper and lower dialects. Lower Sorbian has a native speaker population less than 7000, and the majority of those speakers are older. Hence, it is considered an at-risk and highly endangered language. It is also almost certain that any speaker of Lower Sorbian that is computer literate is also a fluent German speaker, thus the amount of online content in Lower Sorbian is minimal. There is enough, however, that ChatGPT is capable of producing Lower Sorbian content.

The following is a short story in Lower Sorbian generated by ChatGPT (Free) on May 17th, 2024:

W jadnej małej wjesce jo žyła mała źiwka mjenja Mak. Mak jo rada w
pśiroźe była a pśez łuki ze chłopkami běgała. Jadyn źeń, gaž jo słónce so
skłanjało a źěśi su se w suknjach a pśinach wjeseliły, Mak jo zwuk wót doliny słyšyła. Był to zwuk starego chórka, kótaryž Mak až dotychčas njeznawała.

Załožocy radźiłstwo, zwuk njošy do lesa. Gaž pśišła do zapomnionego městna, Mak jo wiźeła staru drjewjanu chórku. Ona jo stała tam z zapomnjonymi spomnjeśimi. Mak jo so až k tomu chórkowu přiblěžila a gaž jo se skłoniła a do wótwrjeteju duri došła, wiźeła jo źědnego źony w šedrym płasći, kótaryž jo na starym pijanźe hrał.

Źona jo se na Mak glědała a se zasajźiła. "Witaj, Mak! Ja som Hanka, starodawna grajka."

Mak jo so pśituliła a pšašo: "Co tu hrajoš?"

Hanka jo se smějo a rěcy: "To jo hudźba pśirody. Kuždy ton jo wót dusy wětrje, źo se do gałuzow a łopjaw zaplěta."

Mak jo z rady glědała a zaspyśeła se až do chórki. "Móžoš mi wučyś?"

Hanka kiwki swójoju głowu a zachinaše. Ona jo grała, Mak jo wučowała, a pśez tjasne cesty a bóle, hudźba jo se jej wjeski odkłucała.

Mak jo se radźiła. Ona jo se hraś wučowała, a ona jo wiźeła, kak hudźba w pśiroźe słychaś.

A tak, w tej małej wjesce, hudźba Maka a Hanki dale hrań se změniła a
wšyścy luźe su se pśipojeli, aby so w zapomnjonych spomnjeśach
jednotały.
    

On the surface, this seems like a very good thing—a large language model producing content in a very small and endangered language. In many ways, it is a very good thing, but there are elements of danger to it that may not be so obvious.

There exist very few people to actually check this story and confirm its correctness as a short story written in Lower Sorbian. I, as a speaker of Polish, can verify that it reads like a Lechitic language; the overall structure and most of the vocabulary is very familiar and recognizable to me. I can say that it is seemingly Lechitic, and certainly not Polish. Is it Lower Sorbian? That is difficult to confirm; maybe it is the very closely related language Upper Sorbian. Even in the circumstance that it is Lower Sorbian, there are more questions that exist. These are questions that are not linguistic, but cultural. Is the story one that makes sense? Is it structured in the way of a Lower Sorbian story? Is it a truly Lower Sorbian story, or is it an English story that has been translated into Lower Sorbian? Is it an anachronism that perpetuates harmful stereotypes?

The problem is that it is likely that the story is not a very good example of Lower Sorbian prose. It has been shown that the quality of the output of an LLM degrades quickly when languages that are not in the top training languages (English, French, German, etc.) are used and that cultural contexts outside of these dominant cultures are also lost. When it comes to LRLs and especially ELRLs, this problem will only exacerbate over time. Recent research has shown that LLMs trained on data produced by other LLMs eventually break down and "collapse".

In high resource languages, this is an issue that can be accounted for and avoided; however, it may not be the case for LRLs. In the case of Lower Sorbian, even this paper may now be causing an issue. The above text is now one of a limited number of pieces of Lower Sorbian prose that a language model may be trained on. It is disproportionately meaningful as a training sample when compared to a similar generation in a language such as English. Any problems will cascade forwards, and over time may become part of LLM Lower Sorbian, which no longer accurately represents true Lower Sorbian. Generally, it is very possible that "model collapse" can happen much faster and much more subtly with LRLs than is otherwise the case. In a future filled with LLM usage, this could cause a fabricated language to become the accepted representation of a low-resource language, with very little way to detect or revert this occurrence.