Where Does it Stand?
Consider this a sequel to our earlier article on Generative AI Translation. While Generative AI’s problems arise due to its general application and wide dataset, proponents of machine translation point towards its more focused design as a reason for its translation quality. Many companies agree with this sentiment, as Google, Microsoft, and Amazon all use machine learning models for their translation software. However, there are still some concerns when utilizing these models that companies should consider before throwing all of their resources behind them.
What is Machine Learning?
Machine Learning is another form of artificial intelligence that has many elements that separate it from other AI models, such as Generative AI. The simplest difference is machine learning’s more focused scope, as it is designed with a specific goal in mind, such as translating a work from Spanish to English. Generative AI programs like Chat GPT offer a more generalized output, meaning they can be used for a multitude of tasks, ranging from drafting a work email to finding an eggplant parmesan recipe. This generalized approach, while great for simpler tasks, starts to break down once more and more technical topics are introduced. We discussed this in our article on Generative AI translation, but the long-and-short of it is that the model’s heavy reliance on building its dataset through scraping the internet means that it has difficulty with assigning relevancy to a datapoint and can easily become confused when it doesn’t have a lot of data to draw on, which is common for languages that do not have a large amount of representation on the internet, like Hindi.
Machine Learning bypasses these concerns by not only having a much more focused dataset, but by having a dataset created by human hands, ensuring that all of the information entered into the system is relevant. In a large language model, or LLM, this is usually accomplished by assigning numerical values to all parts of a sentence, from grammar to syntax to word choice and so on. It then analyzes these numerical values and takes note of patterns that arise. So, for example, the model’s creator might create a numerical value of X, which is assigned to all verbs. By analyzing thousands of sentences containing words marked with X, the machine slowly “learns” where to place these words in relation to words marked with Y, Z, etc. When a Machine Learning AI is then tasked with crafting content, it draws on this recognized pattern to craft a new sentence, cross referencing with possibly millions of data points. Because of this more deliberate design philosophy, many companies like Amazon have used Machine Learning methods to craft their translation software.
Limitations: A Question of Data
While Machine Translation has its benefits, it still stumbles over many of the hurdles affecting Generative AI models, albeit less so. The most pressing question that no AI model has yet to solve is how to handle low-resource languages, or languages that have a very limited amount of text easily accessible. A 2021 study titled “Machine Translation in the Covid domain: an English-Irish case study for LoResMT” focused on the difficulties of translating highly technical documents to the Irish language due to the small amount of easily accessible text to draw upon to build a dataset. Despite explosive growth in the development of AI technology over the past three years, we see this problem continuing to crop up. Most recently, a study titled The Task of Post-Editing Machine Translation for the Low-Resource Language, a study published in 2024, discussed the effectiveness of translating from English and Russian (two languages with a high amount of text available) to Kazakh and Uzbek (two languages with a low amount of text available) and reported much of the same information as earlier studies. However, where this study differentiated from earlier ones was in discussing how AI can increase productivity when paired with a knowledgeable translator to serve as an editor. This study ultimately pointed towards AI becoming a tool for a translator to use in the same way as a word processor, and that companies that rely on these are setting themselves up for failure.
There are also several issues that have yet to be fully explored by studies. For example, very little time has gone into exploring how these large language models tackle different dialects, which is a huge issue in the translation industry due to how wildly different dialects under the same language umbrella may diverge. You might not use the same rules for British English that you would for Jamaican English, which is heavily influenced by Jamaican Patois and other creole languages.
Conclusions
Much of the same conclusions of our generative AI article match up with our thoughts on Machine Translation, though it must be said that Machine Translation appears to be a much more fertile area to explore. While we would never recommend Machine Translation over a skilled translator, it would be useful if translators began exploring Machine Translation as a method for checking their work and increasing their productivity.


