Where Does AI Translation Stand?
Since ChatGPT’s release in November of 2022, any discussion of the internet has come to revolve around the use of generative AI. It’s hard to name a single industry that has not been impacted by the AI boom, and translation services are no different. Meta, Google, and Amazon have all launched AI translation initiatives, and while the technology behind them is impressive beyond a doubt, the question of if these models can handle translating intensive technical, legal, and medical documents still remains. On top of these questions, recent literature poses some interesting counterpoints to the tech industry’s claims of AI’s propensity for near infinite growth.
Defining Terms
Before we dive into the nitty-gritty, let’s define three terms that have been in the news a lot.
- Generative AI – An umbrella term for most AI programs that generate original content, such as those made by ChatGPT and Google’s AI. These programs work by scraping through data points (usually sourced through the internet) and using them as a reference to generate a given piece of content. So, let’s say you are using an AI image generator like Midjourney and you prompt it to give you a picture of “A Rottweiler riding a bicycle.” Midjourney then sorts through its library of millions of internet pictures looking for images it believes are related to the prompt, from the very focused like images of Rottweilers and bicycles to the very broad like dogs and vehicles. It then uses all of these images, which might number in the thousands, as reference points for generating a new image.
- Machine Learning – The process by which an AI “learns.” Basically how it comes to recognize a given data point as relevant for a given prompt. Machine learning is used to help generative AI develop their data sets and recognize what data is relevant, but keep in mind that they are not interchangeable terms. Certain AI models, such as Amazon’s translation service, advertise themselves as Machine Learning Models because they do not “generate content,” meaning their use is much more limited but more consistent.
- Large Language Model – A type of generative AI, large language models focus solely on generating text which, of course, includes translation services.
The Main Issue with AI Translation: A Lack of Data
A recent paper titled “No ‘Zero-Shot’ Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model Performance” has made waves in the realm of AI. This paper, written in a joint effort by scientists from the University of Tübingen, University of Cambridge, and the University of Oxford, presented data that raises questions about the tech industry’s claims about the infinite potential of AI. It is important to note that the paper does not focus on AI translation specifically, but on the AI ecosystem as a whole. One of the paper’s major claims is that, while the data sets fueling these AI models are massive, their choice to generate these data sets by scraping the internet is going to lead to certain categories of information being over represented. You want to generate a picture of a cat? Perfect, there’s millions of pictures for the AI to reference. You want to generate a picture of an Indian Elephant? More difficult, but still, there are thousands of pictures out there. You want to generate an image of a Saint-Francis Satyr, an extremely rare species of butterfly? Now the AI has much less data to draw from and inform its decision making. Now, the AI is having to rely more and more on tertiary data that is related to the given subject but is not a direct representation. So, your Saint-Francis Satyr prompt will probably look like a butterfly, but it might lack details important to the species, such as its iconic patterning or ruddy brown coloration.
With this drawback in mind, let’s apply this to large language models and the translation industry. The vast majority of text sampled by large language models comes directly from the internet, and that is important because the internet is primarily written in English. Recent estimates provided by W3 Techs, an online survey company, theorize that 49.7% of all text on the internet is written in English. The next highest language, Spanish, only makes up 5.9% of all internet text. As you go further and further down the list, many popular languages, such as Chinese or Hindi,have striking gaps in their representation (1.2% and less than 0.1%, respectively). These gaps, of course, are going to have a direct effect on any AI’s ability to translate that language.
So, the main issue then is a problem of data. Many languages just do not have easily accessible, easily scrapable data for these language models to access. The easy answer to this, then, is that given time, these models will catch up as more and more of the information on the internet is transcribed into different languages. Internet coverage continues to expand across the globe and the percentage of the internet written in English decreases every year, with six of the next ten most common languages seeing increases in their use. So, it’s just a matter of time, right?
Flatlining: The Limits of Data
The other broad point the study makes is the limits of adding more and more data into these models. Across the spectrum, we see the efficacy of these models slowly flatline as more and more data enters the system, with some models even dipping in quality as their databases increase in size. Why these programs are stagnating is not totally clear. A possible reason for this dip in quality is a concept called “Model Collapse,” though it is also colloquially referred to as “AI Cannibalism.” Basically, as more and more content on the internet is created by AI, that content feeds back into the system, becoming just another machine learning data point. Some theorize poorly generated AI content can still enter this system, informing the next generation of AI content, which then itself gets fed back into the system. This means that these mistakes continue on and compound, reducing the overall quality of the content generated.
Model collapse was highlighted as a core issue with Large Language Models in a recent study titled “The Curious Decline of Linguistic Diversity: Training Language Models on Synthetic Text.” In the study, the team found that influxes of synthetic, or AI generated, content injected into a machine learning dataset caused a decrease in the quality and diversity of the outputted content. What they witnessed was a general trend of produced content becoming “samey” and developing an artificial cadence to it, with certain mistakes persisting across models. The risks associated with model collapse should become readily apparent when paired with the fact that estimates already point to around 10% of the total content on the internet being AI generated. This content is already starting to filter back into the system, and while its effect has yet to be studied on a large scale, most research points towards it being a net negative on the machine learning ecosystem.
None of these issues disappear when we focus on just AI translation. If anything, the lower amount of data available for these languages makes them particularly vulnerable to model collapse, as synthetic content written in Chinese will have more weight in these large language models than comparable synthetic content written in English due to its limited supply.
Conclusions
It is not difficult to see how the fervor surrounding these large language models is well earned. The excitement backing AI programs like ChatGPT was proportional to the speed at which this technology was developing. Every month, they outdid themselves, constantly improving, to the point where their momentum seemed infinite. Now, though, we are reaching a plateauing period, where the path forward seems unclear.
At the moment, the issues with AI in producing translated content are too overwhelming to support its use. There will surely be more and more developments as time progresses, developments that we hope can be highlighted in the next three months, but until then, we recommend working with a trusted language services agency if you are in need of translation.


