pctechguide.com

  • Home
  • Guides
  • Tutorials
  • Articles
  • Reviews
  • Glossary
  • Contact

How neural machine translation systems based on AI work

Just a few weeks ago, Meta presented an artificial intelligence model capable of translating into 200 languages. The bet on this technology, which has the name ‘No Language Left Behind’ (NLLB-200), is part of a project developed by Mark Zuckerberg’s company to boost its bet on the metaverse.

Almost all the technology giants, with the exception of Apple and Google, are undertaking projects to position themselves in this new virtual universe that is in full swing. But there are other, more modest companies, some of them local, that have long since initiated research efforts in this field. For some years now, the company Incyta and the GRIAL research group of the Arts and Humanities Department of the Universitat Oberta de Catalunya (UOC) have been collaborating on a series of research and technology transfer projects related to neural machine translation. The objective of the research is to develop neural machine translation systems to be integrated into the workflow of the company Incyta.

This Barcelona-based language services company has been using machine translation systems for years to carry out post-editing. This workflow based on machine translation plus post-editing makes it possible to offer a more efficient and economical translation service, while maintaining the same level of quality, to its wide range of clients: the written press, publishing houses, public administration, universities, etc.

Until a few years ago, machine translation systems offered sufficient quality only for similar language pairs, such as Spanish-Catalan or Spanish-French. On the other hand, for slightly more distant languages, such as Spanish-English, for example, the quality of machine translation was not sufficient. It was more efficient to translate the document manually from scratch.

The emergence of today’s neural machine translation systems has made it possible to obtain outstanding quality even for very distant language pairs, such as Chinese-Spanish, for example. The appearance of these systems has constituted a true revolution in the world of professional translation, since they open the door to applying the most post-edition machine translation flow to most translation jobs.

Rule-based machine translation and corpus-based machine translation

But to understand what this technological revolution is all about, it is worth remembering the two main paradigms of machine translation: rule-based machine translation and corpus-based machine translation. In the first paradigm, the rule-based paradigm, machine translation systems are developed by computer engineers and linguists who write programs, dictionaries and rules to translate a sentence in a source language into a sentence in the target language.

The development of these systems usually involves many months of work by teams of several people. Among the rule-based systems, syntactic transfer systems can be highlighted. In these systems, the sentence in the source language is syntactically parsed to automatically obtain a parse tree. This parse tree, which can be deep or shallow, is transferred to an equivalent tree in the target language using a set of rules.

Once this syntactic tree is obtained in the target language, the words are translated using bilingual dictionaries and the translated words are inflected to obtain a correct sentence in the target language. This paradigm has worked very well for similar languages that have quite similar syntactic structures. There are excellent systems using this methodology that are still in use for similar language pairs such as Spanish and Catalan.

In the second paradigm, corpus-based systems, systems are not developed, but trained. That is, the systems learn to translate from texts in the source language and in the target language. Parallel corpora, i.e., sets of segments or sentences in one language with their translation equivalents in another language, are normally used to train these systems.

Chronology of machine translation


The first corpus-based systems are statistical machine translation systems, which burst onto the market around 2005. These systems are based on the calculation of two probabilities: the probability that a given sentence in the target language is the translation of a sentence in the source language; and the probability that a given sentence in the target language is a correct sentence in that language. The first probability can be calculated from the statistics obtained from the parallel corpus; while the second probability is calculated from the statistics obtained in a monolingual corpus of the target language. This monolingual corpus can be obtained from the target language part of the parallel corpus.

Filed Under: Articles

Latest Articles

ISO 9660 Data Format for CDs, CD-ROMs, CD-Rs and CD-RWs

ISO 9660 is a data format designed by the International Standards Organisation in 1984. It's the accepted cross-platform protocol for filenames and directory structures. Filenames are restricted to uppercase letters, the digits 0 to 9 and the underscore character, _. Nothing else is … [Read More...]

Slow Computer

5 Tips to Cure PC Slowness

Today's PCs are designed to be fast and agile. No one wants a computer that takes ages to boot up and even longer to load web pages. This isn't the 90s anymore where slow computers were more common. But, sometimes a computer can get gummed up and may need a little TLC to get it back to its old self. … [Read More...]

Microsoft Debacle Emphasizes Importance of Quality Control in Software

Microsoft set off a chain of criticism when it rolled out its Windows update in 2018. This was a huge debacle that created a lot of frustration for countless customers, which reflected poorly on Microsoft’s brand image. Microsoft responded by assuring customers it would make quality control a … [Read More...]

Top Taplio Alternatives in 2025 : Why MagicPost Leads for LinkedIn Posting ?

LinkedIn has become a strong platform for professionals, creators, and businesses to establish authority, grow networks, and elicit engagement. Simple … [Read More...]

Shocking Cybercrime Statistics for 2025

People all over the world are becoming more concerned about cybercrime than ever. We have recently collected some statistics on this topic and … [Read More...]

Gaming Laptop Security Guide: Protecting Your High-End Hardware Investment in 2025

Since Jacob took over PC Tech Guide, we’ve looked at how tech intersects with personal well-being and digital safety. Gaming laptops are now … [Read More...]

20 Cool Creative Commons Photographs About the Future of AI

AI technology is starting to have a huge impact on our lives. The market value for AI is estimated to have been worth $279.22 billion in 2024 and it … [Read More...]

13 Impressive Stats on the Future of AI

AI technology is starting to become much more important in our everyday lives. Many businesses are using it as well. While he has created a lot of … [Read More...]

Graphic Designers on Reddit Share their Views of AI

There are clearly a lot of positive things about AI. However, it is not a good thing for everyone. One of the things that many people are worried … [Read More...]

Guides

  • Computer Communications
  • Mobile Computing
  • PC Components
  • PC Data Storage
  • PC Input-Output
  • PC Multimedia
  • Processors (CPUs)

Recent Posts

Utilizing the Where Command with Ruby

We recently covered an article on the major benefits of Ruby on Rails. You may want to read a little bit more about it if you are still trying to get … [Read More...]

Advertising on PCTG

PCTechGuide.com is one of the oldest and most respected tech websites and is maintained by Brain Box Consultants LLC. Thousands of universities … [Read More...]

LP to CD Labeling

Although a CD jewel case won't allow you to reproduce the complete liner notes of your original LP, it does give you sufficient space to do a lot … [Read More...]

[footer_backtotop]

Copyright © 2026 About | Privacy | Contact Information | Wrtie For Us | Disclaimer | Copyright License | Authors