Knowledge & learning

Context RL

Reinforcement learning as a service.

The loop that converts captured traces, rubric grades, and expert feedback into measurable improvement: better routing, distilled models you own, and releases that ship only when the evals say they should.

Context RL
Example workflow
Verified traces
Train candidate
Held-out eval
Release gate
Held-out evaluation · example Gate passed
Policy adherence8294
Evidence quality7891
Exception handling7488
Release comparison

Baseline

78

overall score

Candidate

91

overall score

All configured thresholds cleared

The release record keeps the dataset, grader, model, and rubric versions together.

How it works

From request to a reviewable result.

Each run keeps the request, actions, evidence, and final decision together.

  1. 01 · Curate

    Build from verified traces.

    Context selects rubric-passing runs, removes low-confidence labels, and keeps a held-out set for release decisions.

    Output · Versioned dataset

  2. 02 · Improve

    Train or optimize the candidate.

    Run distillation, routing, or policy optimization against the approved training split and grader.

    Output · Candidate version

  3. 03 · Gate

    Compare before release.

    The candidate is evaluated against the held-out suite and promoted only when it clears the configured thresholds.

    Output · Release decision

What you can inspect

Capability is only useful when it leaves evidence.

Verified training data

Rubric-passing traces harvested from production, with rejection sampling on grader confidence.

Recorded with the runTrace and label lineage

Grader alignment first

Judges are validated against expert consensus before any optimization runs against them.

Recorded with the runAgreement by rubric

Distillation to models you own

Open-weight students trained on your verified work, deployed behind Context Inference.

Recorded with the runDataset and model version

Eval-gated releases

Candidates ship only after clearing the held-out thresholds your team defines, with regressions visible by rubric and cohort.

Recorded with the runBaseline comparison

Deployment boundaries

Control is part of the product.

Context RL can run as a standalone surface, while its traces, evals, and policy decisions remain part of the same Context system.

A gate reflects its suite

Held-out evals expose measured regressions; they do not guarantee performance on cases the suite does not represent.

Experts define success

Context automates the loop around rubrics, but domain owners still approve graders, thresholds, and releases.

See how it works with Evals

Talk to us.
Bring a workflow your team runs today and see it run in your environment.