| LLM training & fine tuning |
Curated monolingual, instruction, domain and specialist text with controlled cleaning and metadata. |
JSONL, Parquet, structured text or agreed model format. |
MSA plus the dialect, register and domain evidence required by the intended users. |
| Arabic AI evaluation |
Human reviewed prompts, responses, documents or conversations with expected outcomes and scoring criteria. |
Gold sets, scorecards, adjudicated labels and regression sets. |
Country, dialect, register, domain, channel and risk. |
| ASR & voice AI |
Natural or prompted speech with transcription, acoustic metadata and speaker information. |
Audio, transcript and structured metadata in the agreed format. |
Gulf, Egyptian, Levantine, Maghrebi and defined local varieties. |
| Conversational AI |
Multi turn interactions, intent, entities, escalation events and realistic user language. |
Dialogue structures, JSONL and task specific labels. |
Informal Arabic, local terminology and code switching where representative of real users. |
| RAG & enterprise search |
Trusted institutional, professional and domain content prepared for retrieval and knowledge grounding. |
Clean documents, chunks, metadata and source references. |
Formal Arabic, technical terminology and local institutional usage. |
| Computer vision & multimodal AI |
Arabic visual environments, OCR, video, speech and synchronized annotation according to the model task. |
Images or video with text, events, objects, segments or other project specific labels. |
Region specific signage, environments, objects and usage contexts. |