Model Alignment & Human Feedback
Rubric & Reward Data for LLM Post Training and Evaluation
Rubric and reward data is the human authored material that tells a model, criterion by criterion, what a good response contains, what it must avoid and how much each element should influence the reward. Pangeanic designs instance specific rubrics, preference data, verifiable task sets and grader calibration data in the languages your model will actually serve, so the reward signal remains intelligible to the experts responsible for the task.
The reward problem
Your model will learn exactly what your reward pays for
Reinforcement learning is ruthless about incentives. Whatever the reward signal pays for, the policy will eventually explore. For tasks such as mathematics and code, deterministic verification can provide a comparatively clean signal: an answer matches, a constraint is satisfied or a test passes.
The Tülu 3 work formalized this approach as Reinforcement Learning with Verifiable Rewards (RLVR), replacing a learned reward model with verification functions for suitable tasks. DeepSeek R1 took a related decision for reasoning: its authors explicitly avoided neural reward models for reasoning tasks because they had observed susceptibility to reward hacking during large scale reinforcement learning.
Most enterprise tasks offer no such convenient answer key. A clinical explanation, legal analysis, customer response in Brazilian Portuguese or policy sensitive refusal can be good or bad across several independent dimensions. Accuracy, completeness, relevance, safety, terminology, register and concision can all influence whether the response is useful.
Pairwise preference data provides one solution, but a simple chosen versus rejected label can conceal why one answer was preferred. Rubrics make more of that judgment explicit by decomposing quality into criteria that can be inspected and scored separately.
A rubric makes the reward more interpretable and auditable, while moving much of the quality problem toward rubric specification, expert judgment and grader calibration.
Of course, a rubric remains a proxy. A badly written criterion can be exploited just as efficiently as a badly designed reward model. That is the data problem Pangeanic addresses.
Published evidence
Why AI post training is moving toward explicit expert criteria
Public research already shows how expert rubrics, verifiable rewards and multilingual preference data are changing model training and evaluation.
OPENAI · HEALTHBENCH · 2025
48,562 expert written criteria
HealthBench evaluates 5,000 health conversations against 48,562 unique rubric criteria. The benchmark was created with 262 physicians who have practiced across 60 countries and are collectively proficient in 49 languages.
SCALE AI · RUBRICS AS REWARDS · ICLR 2026
Up to 31% relative gain on HealthBench
Using checklist style rubrics as on policy reward signals, the strongest Rubrics as Rewards variant reported improvements of up to 31% on HealthBench and 7% on GPQA Diamond relative to LLM as judge baselines using direct Likert rewards.
COHERE · MULTILINGUAL PREFERENCE · 2024
Multilingual data improves cross language transfer
Preference optimization trained only on English data achieved a 46.3% win rate on unseen languages. Including 5 training languages increased it to 54.9%, providing direct evidence that multilingual alignment data can improve behavior beyond English.
Decision framework
Which reward signal does your model need?
Most post training programs combine several signals. The useful distinction is whether correctness can be checked mechanically, whether quality is multidimensional and how much expert interpretation the task requires.
| Reward signal | Best for | Data required | Typical failure | Pangeanic delivery |
|---|---|---|---|---|
| Verifiable reward (RLVR) | Math, code, extraction, structured outputs and constraint following | Prompts with checkable reference answers and verifier specifications | The verifier may cover only a narrow definition of success | Expert authored tasks, reference answers, tolerances and verification rules |
| Preference data (RLHF, DPO) | Tone, helpfulness, style and holistic quality | Prompts, multiple responses, rankings or chosen and rejected labels, preferably with rationales | Length, formatting and annotator preferences can become unintended proxies for quality | Rationale backed rankings, adjudication and controls for known preference biases |
| Rubric based reward | Open ended expert tasks in medical, legal, financial, technical and public sector domains | Instance specific criteria, weights, penalties and reference guidance | Under specified criteria, overlapping requirements and grader error | Expert written rubrics with positive, negative and penalty criteria plus grader calibration |
| Process supervision | Long reasoning tasks where the final answer conceals where reasoning failed | Step level judgments on model reasoning or intermediate outputs | High annotation cost and inconsistent segmentation of reasoning steps | Step level expert annotation connected to our expert reasoning workflow |
What we build
Rubric and reward datasets for training and evaluation
Each asset can be delivered independently or combined into a post training data program managed through PECAT, our human governed annotation, review and data operations platform.
01
Instance specific rubrics
Prompt level criteria written by domain experts, including essential requirements, desirable content, prohibited behavior, penalty criteria, weights and reference guidance.
02
Grader calibration sets
Model responses scored criterion by criterion by qualified human reviewers, allowing teams to measure where an LLM judge agrees or diverges from expert judgment.
03
Preference data with rationales
Pairwise or listwise rankings with written reasons, tie handling and adjudication, designed to expose why one response is preferred rather than preserving only the final label.
04
Verifiable task sets
Problems with unambiguous reference answers, tolerances and verifier specifications for RLVR, connected to our Expert Reasoning Data workflow.
05
Reward hacking and grader audits
Adversarial outputs designed to expose criteria that can be satisfied literally while producing incorrect, irrelevant, unsafe or strategically padded responses. Confirmed failures become revised criteria and reusable regression cases.
06
Multilingual parity rubrics
Rubrics authored or substantively adapted by target language experts, with language specific grader calibration and parity checks designed to test whether equivalent behavior receives equivalent reward across languages and regional variants.
Rubric design
Five checks we apply to rubric criteria
A rubric becomes part of the reward mechanism as soon as it enters a training loop. The criteria therefore need to survive both human interpretation and optimization pressure.
1. Atomic. One criterion should represent one judgment. A requirement such as "states the dosage and explains the interaction risk" should normally be separated so partial satisfaction is visible.
2. Self contained. The evidence required for judgment should be available from the task, response and permitted reference material rather than from an evaluator guessing the author's intention.
3. Weighted. Importance is expressed explicitly where the training or evaluation design requires it, including penalties for critical errors or prohibited behavior.
4. Language anchored. Criteria are authored or substantively adapted by specialists in the target language and variant, accounting for terminology, professional register, institutions and local context.
5. Grader tested. Agreement is measured across human reviewers and, where applicable, against the model judge. Criteria producing unacceptable disagreement are clarified, rewritten or removed before delivery.
These checks complement the CURVD principles used in our verifiable reasoning workflows. CURVD makes answers easier to verify; these five checks make judgments easier to inspect.
The multilingual reward gap
A translated rubric can reward the wrong answer, fluently
Most public preference and rubric resources remain heavily English centered. When another language is required, machine translation can produce superficially multilingual data while leaving the underlying judgment anchored to English institutions, terminology, professional conventions and expectations.
A legal concept may have a different scope. A clinical warning may be technically present but pragmatically inappropriate. A public administration process may exist only within a particular jurisdiction. A customer response may satisfy every translated criterion while sounding foreign to the people expected to trust it.
The reward model can then become very efficient at paying for an answer that is correct under one institutional culture and wrong under another.
Picture your team preparing a Spanish language benefits assistant for a public administration. The rubric originated in English, the judge agrees with it and the reward curve climbs beautifully. Then the first citizen asks about a procedure governed by regional law. The model follows the wrong institutional logic perfectly because that is exactly what the reward specification taught it to do.
Research on multilingual preference optimization already provides evidence that alignment behavior changes when the training language mix changes. Pangeanic therefore treats translation as one possible input to rubric development, never as proof that the reward specification is valid in the target language.
Our multilingual workflows combine target language expertise, domain review, local terminology and grader calibration so that language differences become part of the reward design rather than an afterthought.
Delivery format
Model ready files your trainer can consume directly
Rubric and reward data can be delivered in JSONL, JSON, CSV or a client defined schema. Records can include prompts, reference material, criteria, weights, penalties, reviewer decisions, rationales, language metadata and adjudication status.
ILLUSTRATIVE RECORD · JSONL
{
"id": "rr-es-0142",
"language": "es-ES",
"domain": "public-services",
"prompt": "...",
"reference_answer": "...",
"criteria": [
{
"id": "c1",
"text": "Identifica la administración autonómica competente",
"weight": 5,
"type": "essential"
},
{
"id": "c2",
"text": "Indica el plazo de presentación correcto",
"weight": 3,
"type": "important"
},
{
"id": "c3",
"text": "Remite a un procedimiento estatal no aplicable",
"weight": -4,
"type": "penalty"
}
],
"review": {
"status": "adjudicated",
"reviewer_count": 2,
"criterion_scores": {
"c1": 1,
"c2": 1,
"c3": 0
}
}
}
AI Data Operations workflow
From target behavior to a validated reward dataset
Define the behavior
Agree on model tasks, target languages, risk profile, policy boundaries and the training or evaluation method the data must support.
Qualify the experts
Match domain expertise, linguistic competence and independent review requirements to each task and language.
Author the data
Create prompts, reference material, preference judgments or instance specific criteria according to the selected reward design.
Weight and constrain
Define required behavior, penalties, tolerances and other scoring rules without allowing one broad criterion to hide several judgments.
Stress test
Test normal and adversarial responses to identify ambiguous criteria, exploitable wording and missing negative requirements.
Calibrate the grader
Compare human expert judgments with automated grading at criterion, task, domain and language level where applicable.
Deliver with evidence
Ship model ready files together with guidelines, review records, adjudication information and agreed quality reporting.
Refine from training results
New model failures, grader disagreements and reward exploits can become revised criteria, additional tasks and regression cases for subsequent rounds.
Commercial applications
Who buys rubric and reward data?
| Buyer | Objective | Pangeanic delivery |
|---|---|---|
| AI laboratories | Extend post training beyond easily verifiable tasks into expert and open ended domains | Expert rubrics, RLVR task sets, preference data, grader calibration and reward hacking audits across languages |
| Enterprise model teams | Align a fine tuned or small language model to internal policy, terminology and user expectations | Policy derived rubrics, preference data with rationales, proprietary task sets and grader calibration |
| Model evaluation teams | Validate an LLM as judge pipeline before using it for benchmarking or release decisions | Human scored calibration sets, agreement analysis and private evaluation rubrics connected to Evaluation & AI QA |
| Public administrations | Align citizen facing systems to jurisdiction specific procedures, terminology and official languages | In language rubrics, jurisdiction aware criteria, expert review and traceable adjudication |
| Regulated industries | Make model evaluation and reward decisions more inspectable for internal governance and audit | Documented criteria, weights, expert judgments, reviewer evidence and versioned rubric sets |
| Multilingual AI developers | Reduce the gap between strong English behavior and weaker performance across other languages | Native preference data, multilingual parity rubrics and language specific grader calibration, including regional and low resource language programs |
Why Pangeanic
Reward data built inside a full AI Data Operations chain
Rubrics rarely fail in isolation. Weak prompts, ambiguous criteria, reviewer disagreement, language asymmetry and poorly calibrated judges all affect the signal that finally reaches the model. Pangeanic can manage those dependencies as one data operation.
Multilingual by origin
Our background in multilingual language data, machine translation evaluation and Machine Translation Quality Estimation provides practical experience with automatic scoring systems whose reliability changes with language, domain and terminology.
Connected to evaluation and red teaming
Rubrics, evaluation data and multilingual red teaming can share failure taxonomies so that problems discovered during testing become criteria, adversarial cases and regression data for future training rounds.
Private and sovereign workflows
Controlled data operations can support unreleased model outputs, confidential policies and proprietary task material, with data masking and anonymization where projects originate from real conversations or sensitive documents.
RELATED EXPERIENCE · BARCELONA SUPERCOMPUTING CENTER
Our work with the Barcelona Supercomputing Center includes multilingual data annotation, human feedback, reward model related data and LLM evaluation workflows. Read the BSC use case →
FAQ
Questions buyers ask about rubric and reward data
What is rubric and reward data?
Rubric and reward data is human authored material used to evaluate model outputs or generate training rewards. It can include instance specific rubrics with weighted criteria, preference rankings with rationales, verifiable tasks with reference answers and human scored calibration sets for automated graders.
What is the difference between a rubric reward and a reward model?
A learned reward model typically infers a scoring function from human preference data. A rubric based reward makes more of the scoring criteria explicit for each task and asks a human or model grader to evaluate them. This makes the criteria easier to inspect and revise, although reliability still depends on the quality of the rubric and the grader applying it.
When should I use RLVR instead of rubric based rewards?
RLVR is particularly suitable when correctness can be checked mechanically, such as a numerical answer, executable code, structured extraction or formal constraint. Rubric based rewards are useful when quality depends on several qualitative judgments that cannot be reduced to one deterministic test.
Can rubric data be used for evaluation as well as training?
Yes. Rubrics can be used to evaluate model releases as well as to generate reward signals during post training. Training and evaluation datasets should remain appropriately separated so the evaluation does not simply measure criteria the model has already optimized against.
Why not simply translate English rubrics into other languages?
Translation can provide a starting point, but the resulting rubric still needs validation against target language terminology, institutions, professional conventions, register and local expectations. Pangeanic therefore uses target language experts to author or substantively adapt criteria rather than treating raw translation as a finished reward specification.
How do you test whether a rubric can be gamed?
Criteria are tested against normal, borderline and adversarial responses, including outputs designed to satisfy wording literally while remaining incorrect, incomplete, irrelevant or unsafe. Exploitable criteria are clarified, separated, reweighted or removed before delivery.
Can Pangeanic calibrate our LLM as judge?
Yes. Qualified reviewers can score model responses criterion by criterion so agreement and disagreement between the automated grader and human experts can be measured by task, criterion, domain and language.
Which formats can you deliver, and can the project remain private?
Data can be delivered in JSONL, JSON, CSV or a client defined schema together with agreed documentation and quality records. Controlled workflows can also be arranged for unreleased model outputs, confidential prompts, proprietary policies and other restricted project material.
Continue exploring
Where rubric and reward data fits in the alignment stack
Model Alignment & RLHF
Preference ranking, multilingual review, human feedback and evaluation loops for dependable model behavior.
Explore model alignment → Verifiable dataExpert Reasoning Data
Expert tasks, verified reasoning traces and controlled answers for supervised fine tuning and RLVR.
View reasoning data → Failure discoveryMultilingual AI Red Teaming
Adversarial testing, confirmed failures and regression data across languages, cultures and deployment contexts.
View red teaming → MeasurementEvaluation & AI QA
Benchmark design, multilingual quality assurance and regression testing for measurable model performance.
Explore evaluation → Operational layerAI Data Operations
Governed data, human feedback, evaluation, alignment and quality control connected across the AI lifecycle.
Explore AI Data Operations → PlatformPECAT
Human governed annotation, validation, review and traceable data operations for multilingual and multimodal AI.
Explore PECAT →Sources
- Gunjal, A. et al. Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains . ICLR 2026.
- Arora, R. K. et al. HealthBench: Evaluating Large Language Models Towards Improved Human Health . OpenAI, 2025.
- Lambert, N. et al. Tülu 3: Pushing Frontiers in Open Language Model Post Training . Allen Institute for AI, 2024.
- Dang, J. et al. RLHF Can Speak Many Languages: Unlocking Multilingual Preference Optimization for LLMs . EMNLP 2024.
- DeepSeek AI. DeepSeek R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning . Nature, 2025.
Rubric Data · Reward Data · LLM Post Training
Pay your model for the right behavior, in every language
From expert rubrics and preference data to grader calibration, RLVR tasks and multilingual reward audits, Pangeanic builds reward data that remains inspectable by the people who understand the task.
A Representative Vendor in the 2024 "Market Guide for Data Masking and Synthetic Data"

