Fragmented repositories
Translation memories were distributed across institutions, suppliers, desktops and incompatible server environments.
Led by Pangeanic, NEC TM Data created a central, open and CAT tool independent platform for storing, searching, exchanging and reusing bilingual data generated by European public administrations.
European administrations generate large volumes of bilingual content through legislation, public services, procurement, institutional communication and international cooperation.
Much of that work is stored in translation memories linked to specific tools, contractors, departments or national systems. Valuable bilingual data can therefore become fragmented across incompatible formats and isolated repositories.
NEC TM Data addressed this problem by creating a shared platform where public translation assets could be stored, searched, retrieved and reused independently of a specific CAT tool or provider.
Translation memories were distributed across institutions, suppliers, desktops and incompatible server environments.
Access to bilingual assets could depend on a particular CAT tool, proprietary server or vendor ecosystem.
Translations financed by public institutions were not always available for future reuse or language technology development.
Users could not easily search by language, domain, similarity or provenance across multiple collections.
NEC TM Data created a shared architecture in which national public administrations could store, retrieve and exchange bilingual assets while remaining connected to translation providers and CEF eTranslation.
The platform reduced fragmentation, improved reuse of publicly funded translations and increased the volume of parallel data available for future language technology.
NEC TM Data combined translation memory processing, central indexing, open interfaces and controlled sharing to create reusable public language infrastructure.
Recover bilingual assets from public administrations, translation suppliers and existing institutional repositories.
Convert heterogeneous translation memory files and metadata into structures suitable for common indexing and reuse.
Remove duplicates, invalid segments, incorrect language pairs and low quality material before publication.
Use Elasticsearch based indexing to support rapid search, filtering and similarity matching across collections.
Allow translation tools, institutional systems and language technology services to retrieve relevant bilingual data.
Turn previous translations into assets for professional workflows, machine translation, evaluation and future AI models.
A translation memory contains aligned source and target segments together with linguistic and operational metadata. When cleaned and governed, these assets can support multiple AI and language technology tasks.
High quality aligned segments can be used to train or adapt specialised translation models.
Repeated bilingual patterns help identify domain terminology and preferred institutional language.
Curated translation pairs can support benchmarking and comparison of language models and translation systems.
Similarity matching allows translators and systems to recover previous institutional translations quickly.
Bilingual data can be grouped by legal, administrative, healthcare, tourism or technical subject areas.
Governed parallel data supports smaller models adapted to defined languages, domains and operational requirements.
Pangeanic coordinated the project and contributed the central technology and data operations required to transform distributed translation memories into reusable infrastructure.
Pangeanic led consortium delivery, technical planning, public sector coordination and European reporting.
The project extended ActivaTM into an open central environment for indexing and retrieving bilingual assets.
Bilingual files were extracted, normalised, cleaned, deduplicated and prepared for ingestion.
APIs made it possible for different tools and institutional systems to query the same language assets.
Search infrastructure supported exact and approximate retrieval across large collections.
Core platform components were released publicly to support inspection, reuse and further development.
NEC TM Data converted distributed translation memories into infrastructure that could support professional translation, public administration and future AI development.
A common platform allowed several institutions and users to access bilingual assets through one environment.
Applications and CAT tools could connect through common interfaces rather than proprietary integrations.
Translation memory exchange formats supported portability between institutional and professional environments.
Core software was released openly to encourage reuse and continued development beyond the project.
The consortium combined translation memory technology, multilingual data expertise, professional language services and public administration requirements.
Project coordination, central platform, bilingual data processing, API design and European integration.
Technology partnerEuropean language technology, national translation memory infrastructure and multilingual data integration.
Language services partnerProfessional translation workflows, translation memory assets and practical validation.
Spanish public administration participation and alignment with European digital language infrastructure.
Existing bilingual assets can become governed training, evaluation and retrieval data when they are properly extracted, cleaned and classified.
Recover translation assets generated across departments and make them reusable for future services.
Convert multilingual archives into training and evaluation datasets for specialised language models.
Consolidate translation memories across suppliers, tools and business units inside a common data layer.
Access clean parallel corpora for adaptation, domain training, benchmarking and multilingual alignment.
These references document the action identifier, project duration, central platform and open software produced through NEC TM Data.
European documentation covering NEC TM Data, its action identifier, funding and infrastructure objectives.
View European documentation →Public source code for the open version of the platform developed through the project.
Explore the repository →NEC TM Data within Pangeanic’s wider work in multilingual datasets, translation infrastructure, privacy and sovereign AI.
Explore all projects →NEC TM Data forms part of a broader European trajectory connecting language assets, translation infrastructure, specialised models and human validated datasets.
Review projects covering multilingual datasets, machine translation, privacy, cultural heritage and sovereign AI infrastructure.
View all projects → Translation orchestrationDiscover the secure, open platform created to route public sector translation requests across multiple providers.
Explore iADAATPA and MT Hub → Neural model infrastructureExplore the direct neural translation infrastructure created for the official languages of the European Union.
Explore NTEU → Reusable multilingual datasetsDiscover reusable AI tools, multilingual datasets and human validation workflows for European cultural heritage.
Explore AI4Culture →Pangeanic supports enterprises and public institutions with translation memory extraction, corpus cleaning, alignment, domain classification, evaluation data and multilingual model training assets.