European bilingual data infrastructure

NEC TM Data: turning public translation memories into reusable language assets

Led by Pangeanic, NEC TM Data created a central, open and CAT tool independent platform for storing, searching, exchanging and reusing bilingual data generated by European public administrations.

2017-EU-IA-0149 Connecting Europe Facility action identifier
2018 to 2020 Eighteen month infrastructure and deployment programme
€926,266 CEF funding for reusable European language data
4 partners Language technology companies and public administration
Pangeanic led Project coordination and central platform development
The language data problem

Publicly funded translations often remained fragmented, inaccessible and difficult to reuse

European administrations generate large volumes of bilingual content through legislation, public services, procurement, institutional communication and international cooperation.

Much of that work is stored in translation memories linked to specific tools, contractors, departments or national systems. Valuable bilingual data can therefore become fragmented across incompatible formats and isolated repositories.

NEC TM Data addressed this problem by creating a shared platform where public translation assets could be stored, searched, retrieved and reused independently of a specific CAT tool or provider.

01

Fragmented repositories

Translation memories were distributed across institutions, suppliers, desktops and incompatible server environments.

02

Tool dependence

Access to bilingual assets could depend on a particular CAT tool, proprietary server or vendor ecosystem.

03

Lost public value

Translations financed by public institutions were not always available for future reuse or language technology development.

04

Limited discoverability

Users could not easily search by language, domain, similarity or provenance across multiple collections.

Shared European language infrastructure

A central platform connecting national translation memories and European services

NEC TM Data created a shared architecture in which national public administrations could store, retrieve and exchange bilingual assets while remaining connected to translation providers and CEF eTranslation.

The platform reduced fragmentation, improved reuse of publicly funded translations and increased the volume of parallel data available for future language technology.

NEC TM Data platform connecting European public administrations, translation service providers and CEF eTranslation
Original NEC TM Data platform architecture showing national translation memory sharing and connectivity with CEF eTranslation.
The NEC TM Data approach

Recover bilingual assets, normalise them and make them searchable through a common platform

NEC TM Data combined translation memory processing, central indexing, open interfaces and controlled sharing to create reusable public language infrastructure.

01

Collect translation memories

Recover bilingual assets from public administrations, translation suppliers and existing institutional repositories.

02

Normalise formats

Convert heterogeneous translation memory files and metadata into structures suitable for common indexing and reuse.

03

Clean bilingual content

Remove duplicates, invalid segments, incorrect language pairs and low quality material before publication.

04

Index centrally

Use Elasticsearch based indexing to support rapid search, filtering and similarity matching across collections.

05

Expose through APIs

Allow translation tools, institutional systems and language technology services to retrieve relevant bilingual data.

06

Reuse as infrastructure

Turn previous translations into assets for professional workflows, machine translation, evaluation and future AI models.

Reusable bilingual assets

Translation memories are structured parallel datasets

A translation memory contains aligned source and target segments together with linguistic and operational metadata. When cleaned and governed, these assets can support multiple AI and language technology tasks.

01

Machine translation training

High quality aligned segments can be used to train or adapt specialised translation models.

02

Terminology extraction

Repeated bilingual patterns help identify domain terminology and preferred institutional language.

03

Evaluation datasets

Curated translation pairs can support benchmarking and comparison of language models and translation systems.

04

Retrieval and search

Similarity matching allows translators and systems to recover previous institutional translations quickly.

05

Domain adaptation

Bilingual data can be grouped by legal, administrative, healthcare, tourism or technical subject areas.

06

Task specific language models

Governed parallel data supports smaller models adapted to defined languages, domains and operational requirements.

Pangeanic in NEC TM Data

Consortium leadership, platform engineering and multilingual data processing

Pangeanic coordinated the project and contributed the central technology and data operations required to transform distributed translation memories into reusable infrastructure.

LEAD

Project coordination

Pangeanic led consortium delivery, technical planning, public sector coordination and European reporting.

CORE

Central platform

The project extended ActivaTM into an open central environment for indexing and retrieving bilingual assets.

DATA

Translation memory processing

Bilingual files were extracted, normalised, cleaned, deduplicated and prepared for ingestion.

API

Open access interfaces

APIs made it possible for different tools and institutional systems to query the same language assets.

SEARCH

Fuzzy matching and search

Search infrastructure supported exact and approximate retrieval across large collections.

OPEN

Open source release

Core platform components were released publicly to support inspection, reuse and further development.

Project results

A common foundation for storing and reusing European bilingual data

NEC TM Data converted distributed translation memories into infrastructure that could support professional translation, public administration and future AI development.

API

Tool independent access

Applications and CAT tools could connect through common interfaces rather than proprietary integrations.

TMX

Standards based exchange

Translation memory exchange formats supported portability between institutional and professional environments.

Open

Reusable software

Core software was released openly to encourage reuse and continued development beyond the project.

European consortium

Language technology providers and public sector coordination

The consortium combined translation memory technology, multilingual data expertise, professional language services and public administration requirements.

From public translation memory to Data for AI

What NEC TM Data demonstrates for enterprises and public institutions

Existing bilingual assets can become governed training, evaluation and retrieval data when they are properly extracted, cleaned and classified.

01

For public administrations

Recover translation assets generated across departments and make them reusable for future services.

02

For enterprise AI teams

Convert multilingual archives into training and evaluation datasets for specialised language models.

03

For translation technology teams

Consolidate translation memories across suppliers, tools and business units inside a common data layer.

04

For model developers

Access clean parallel corpora for adaptation, domain training, benchmarking and multilingual alignment.

Multilingual data for AI

Turn your translation memories and multilingual archives into governed AI assets

Pangeanic supports enterprises and public institutions with translation memory extraction, corpus cleaning, alignment, domain classification, evaluation data and multilingual model training assets.