Training the Next Generation of Translators and Interpreters

Translation and interpretation advancements for existing AI models continue to be on the forefront of desires for many large language model developers, and for good reason. As we have touched on extensively in past articles, many current AI models struggle with translation projects depending on the target language due to a lack of readily available training data stemming from said language lacking in content that can be scraped off the internet. For example, despite a language like Hindi being one of the most commonly spoken languages in the world, LLM’s struggle to translate it due to only 0.04% of the internet being written in it. This flaw is further compounded when the project requires understanding and translating technical jargon, like translating a pre-surgical screening report. Additionally, while developments in AI translation have been slowly pushed through, similar developments in AI interpretation are nearly non-existent, due to the added complexities of intonation, body language, and regional speech differences.

As AI developers pivot to sourcing their translation data from different locales, though, another issue has arisen; that being that many professional translators are both wary and sceptical of AI translation’s capabilities. At 7C Lingo, we understand the inherent reservations many professional translators would hold for this technology, fearing that one day this tool will outright replace all humans in the field. The University of Cambridge has also come to recognize this conflict, and has chosen to emphasize how translators can use AI programs not only to finish projects, but also develop their own skills. With these principles in mind, the university developed AI Conversation Trainer, an active AI assistant designed specifically to coach would-be multilinguists through their journey towards learning another language.

What is Cambridge AI Conversation Trainer?

This new tool from the University of Cambridge is designed not to outright teach a secondary language, but instead analyze a student’s progress towards language mastery in conjunction with a skilled teacher. The way the software works is that the student records their voice in another language and the AI then analyzes it for issues, giving recommendations on where the student can improve their pronunciation. The program goes as far as to break down certain issues within pronunciation, and will highlight individual syllables within sentences that need to have their pronunciation adjusted.

The most robust version of this tool is Cambridge AI Speech Tutor, a version dedicated specifically for instructing users on speaking English. This version includes the ability to switch the tool to analyze various accents, and includes a percent score to help users gauge their progress towards language mastery.

While the baseline Speech Tutor software lacks some of the robustness of its English counterpart, it does include many more languages to work with. At the moment, the model supports English, French, Italian, Norwegian, Spanish, German, Russian, Ukrainian, Polish, and Romanian. Additionally, the bass model can also switch between spoken or written words, allowing for a person to practice all aspects of a language before deploying it. If a user finds that their proficiency level has increased, they can also adjust the difficulty of the practice questions being prompted, with the difficulty being based on the Common European Framework of Reference for Languages, a standard international framework for gauging one’s language proficiency.

What are the Issues?

While the software is certainly impressive, a few issues need to be addressed before it can receive a recommendation. Obviously, the language selection is fairly limited compared to other AI language options available. The languages are also exclusively European, with the majority of them having training data readily available for teaching and training purposes. Cambridge has not indicated if they plan on expanding the language catalog anytime soon, which would be a shame, as it cuts out many languages that are in desperate need of solid translation learning tools. The tool also makes no distinction based on dialect, which could lead to confusion during instruction depending on the dialect of your teacher.

It is also unclear as to what datasets this model has been trained on to ensure accuracy in translation. While it does make mention of Google, it does not specify which dataset it is derived from, as Google has recently made publicly available several data pools for translation and interpretation purposes.

Without further knowledge on how this program was trained and its efficacy compared to other programs in the space, it is difficult to recommend. However, the promise of a tool that works with the user in order to build a greater understanding of the language is commendable, and we at 7C Lingo hope that this is a trend other companies will explore in the near future.