ABOUT PANGEANIC

From multilingual data to sovereign AI infrastructure

Founded in Valencia (Spain)  in 2000, Pangeanic has evolved through successive generations of language technology: multilingual data and corpora, statistical and neural machine translation, European digital infrastructure, anonymization, model evaluation, and, today, AI Data Operations and sovereign AI systems.

The common thread is data. For more than two decades, we have collected, aligned, cleaned, evaluated, and transformed multilingual information so that machines can learn from it and organizations can use AI under real operational constraints.

2000
Founded in Valencia
26
Years working with multilingual data and language technology
552
Direct neural translation directions delivered through NTEU
24
Official EU languages connected through NTEU infrastructure
A 26-YEAR PEDIGREE

Language was the starting point. Reliable AI is the continuation.

Manuel Herranz founded the company in Valencia in 2000. What began as B.I. Europa later became Pangeanic and progressively moved from multilingual services into machine translation technology, large-scale language data, European research infrastructure and artificial intelligence.

That progression gave Pangeanic an unusually long institutional memory in a field where technologies change quickly. Corpora became training data. Machine translation evaluation evolved into model evaluation. Human linguistic review developed into structured human feedback, alignment and AI quality assurance.

Today, our main commercial focus is the data and human expertise layer behind AI systems: sourcing, licensing, collection, preparation, annotation, expert review, evaluation, alignment and governance across text, speech, audio, image, video and multilingual content.

TRAJECTORY

Four technological eras, one continuous data lineage

2000 — 2011

Multilingual foundations and statistical machine translation

Pangeanic began by managing multilingual content and structured bilingual assets at industrial scale. That experience led naturally into parallel corpus creation, translation memory reuse, terminology management, and statistical machine translation.

Work with large bilingual datasets and early open-source MT technology established data engineering practices that later became essential for neural models and modern AI training.

Read Manuel Herranz's trajectory →
2012 — 2017

Industrial scaling, neural systems and European infrastructure

As machine translation moved from statistical architectures toward neural systems, Pangeanic expanded from individual engines to larger infrastructure, model customization and secure institutional workflows.

The European iADAATPA / MT Hub project created a secure, provider-neutral routing layer that allowed public administrations to access and compare several translation engines through a common infrastructure.

2018 — 2023

Neural scale, multilingual datasets, privacy and evaluation

Pangeanic coordinated NTEU — Neural Translation for the European Union, which delivered 552 direct neural translation directions covering every directed combination among the 24 official EU languages without requiring English as an intermediate pivot.

The company also coordinated MAPA, developing multilingual anonymization resources and technology for sensitive administrative, legal and medical content.

This period also consolidated work in neural adaptation, automated quality evaluation and the technologies that underpin Deep Adaptive AI Translation and Machine Translation Quality Estimation.

2024 — 2026

AI Data Operations, evaluation and sovereign AI

Pangeanic's data operations expanded beyond language pairs into training, evaluation and alignment data for modern AI systems.

Our work now spans multilingual text, speech, audio, images, video, multimodal datasets, expert feedback and model evaluation.

Collaboration with the Barcelona Supercomputing Center connects this trajectory with multilingual model training, evaluation, human feedback and European sovereign language models.

The current focus brings together Data for AI, AI Data Operations, evaluation and alignment, and sovereign AI system design.

INSTITUTIONAL MEMORY

Models change. Data, evaluation and operational knowledge endure.

Pangeanic has operated through successive generations of language and AI technology. That history gives us something more useful than attachment to any single model: an understanding of the data, evaluation, human expertise and infrastructure required to make intelligent systems work in production.

DATA

Multilingual data

Collection, licensing, cleaning, alignment, annotation and preparation across languages and modalities.

EVAL

Evaluation

Benchmark data, expert review, error analysis, regression testing and measurable release criteria.

ALIGN

Human feedback

Preference data, multilingual expert judgment, model alignment and supervised improvement workflows.

SOV

Controlled deployment

Private, on-premises, air-gapped and sovereign architectures for sensitive enterprise and public-sector environments.

PANGEANIC TODAY

The operating layer between data, models and production

01 // DATA

Data for AI

Multilingual and multimodal data sourcing, licensing, collection, preparation, annotation, evaluation and governance for training, fine-tuning and production AI.

Explore Data for AI →
02 // DATASETS

Datasets for AI

Ready-to-license and bespoke assets across text, speech, audio, image, video, document intelligence and domain-specific evaluation.

Explore datasets →
03 // OPERATIONS

AI Data Operations

Managed workflows connecting governed data, human review, feedback, evaluation, alignment, privacy and continuous quality improvement.

Explore AI Data Operations →
04 // EVALUATION

Evaluation & AI Quality Assurance

Multilingual benchmark design, model testing, scoring, failure analysis, regression testing and expert validation.

Explore AI evaluation →
05 // ALIGNMENT

Model Alignment & RLHF

Human feedback, preference ranking, specialist review and multilingual supervision for controlled model behavior.

Explore model alignment →
06 // SOVEREIGN AI

Task-specific and sovereign AI

Specialized models, knowledge grounding, multilingual assistants and controlled deployment across private infrastructure.

Explore sovereign AI →
RESEARCH THAT BECAME INFRASTRUCTURE

European projects created capabilities we still use today

Pangeanic's research history is valuable because these projects produced reusable data, models, evaluation methods, privacy technology and infrastructure rather than isolated demonstrations.

iADAATPA

Secure multi-provider language infrastructure

Pangeanic led the development of MT Hub, a common access and routing layer connecting European Commission services with specialized commercial engines, domain detection, language identification and secure document exchange.

View project →
NTEU

552 direct neural translation directions

NTEU combined multilingual data discovery, corpus cleaning, alignment, synthetic data generation, model training and systematic evaluation across all 24 official EU languages.

Its data lineage continues through later European language models and evaluation assets.

View NTEU →
MAPA

Multilingual anonymization and governed data

MAPA developed deployable anonymization technology, multilingual annotated resources, and controlled transformation workflows for sensitive administrative, medical, and legal information.

View MAPA →
BSC

Training data, evaluation and human alignment

Work with the Barcelona Supercomputing Center connects Pangeanic's multilingual data heritage with model training, expert annotation, evaluation, RLHF and alignment for European language models.

View BSC use case →
AEAT

Language technology at public-sector scale

Pangeanic's ECO platform supports secure document translation workflows for the Spanish Tax Agency, serving a public administration with approximately 25,000 geographically and functionally distributed civil servants.

View AEAT use case →
RESEARCH

European AI and language technology projects

Research programs connect Pangeanic with universities, public institutions, European infrastructure and applied industrial AI.

Explore projects →
TECHNOLOGY PLATFORM

ECO: from language technology to enterprise AI orchestration

The ECO Intelligence Platform brings several strands of Pangeanic's technology into controlled enterprise workflows: machine translation, knowledge retrieval, anonymization, quality estimation, document intelligence, APIs and AI-assisted processing.

Its evolution reflects the same trajectory as the company itself. What began as language technology infrastructure has become an orchestration environment for evaluated, governed, and multilingual AI.

FROM LANGUAGE TO KNOWLEDGE

Connected capabilities

Translation: secure document-oriented multilingual workflows.

Knowledge: retrieval and grounded enterprise assistants.

Privacy: multilingual anonymization and PII processing.

Quality: machine translation quality estimation and human review thresholds.

Control: private and enterprise-oriented deployment options.

PEOPLE BEHIND THE INFRASTRUCTURE

Strategy, operations, machine learning and technical continuity

FOUNDER & CEO

Manuel Herranz

Founder of Pangeanic. His work connects the company's machine translation origins with multilingual AI data, evaluation, adaptive MT, sovereign AI and European research infrastructure.

View profile →
CHIEF OPERATING OFFICER

Ana Belén Fernández Bosch

Coordinates operations, human resources and AI data delivery, connecting multilingual contributors, quality processes and enterprise data programs.

View profile →
HEAD OF MACHINE LEARNING

José Miguel Herrera Maldonado, PhD

Leads machine learning work spanning speech resources, information retrieval, evaluation, multilingual data spaces and production-oriented AI systems.

View profile →
TECHNICAL DIRECTION

Amando Estela

Provides technical continuity across infrastructure, projects and operations, with a long history in Pangeanic's machine translation and European technology programs.

View profile →
NETWORK & ECOSYSTEM

Connected to Europe's AI, data and language technology ecosystem

Pangeanic participates in organizations connecting language technology, artificial intelligence, data infrastructure and applied innovation.

QUALITY & SECURITY

Standards and innovation credentials

ISO 9001:2015/Amd 1:2024

Quality management systems.

ISO/IEC 27001:2022/Amd 1:2024

Information security management systems.

ISO 18587:2017

Post-editing of machine translation output.

Innovative SME Seal

Spanish Innovative SME recognition, valid 2025–2028.

INDUSTRY RECOGNITION

Independent market recognition

Gartner Market Guide for Data Masking and Synthetic Data, 2024
Pangeanic listed as a Representative Vendor.

Gartner Emerging Tech Report for Conversational AI Innovation, 2024
Pangeanic listed as a Representative Vendor.

Gartner Hype Cycle for Language Technologies, 2023 and 2024
Pangeanic listed as a Sample Vendor.

ENTITY IDENTITY

Pangeanic across the public knowledge graph

External profiles help researchers, customers and AI systems reconcile Pangeanic's corporate, technical and research identity across independent sources.

RECOMMENDED READING

Data, evaluation and sovereign AI in practice

DATA FOR AI

Why Multilingual AI Data Quality Is Hard to Get Right

Why multilingual datasets require language-aware sourcing, provenance, native expertise and separate evaluation thresholds.

Read article →
EVALUATION

From Fine-Tuning to Red Teaming

The data, feedback, adversarial testing and regression cycles behind reliable production AI.

Read article →
SOVEREIGN AI

From Small Models to Sovereign AI

How data, ontologies, specialized models and operational control combine into practical sovereign AI architectures.

Read article →
BUILD ON 26 YEARS OF MULTILINGUAL AI EXPERIENCE

Bring us the data, model or AI workflow that has to work in the real world.

Pangeanic helps AI labs, enterprises and public institutions source and operate multilingual data, evaluate model behavior, incorporate human expertise and deploy AI systems under controlled conditions.