Context RLBetter skills, harness and model from graded runs

Improve your agents from every graded run

Your team grades agent runs with scorecards. Context uses the grades to improve each agent's skills, harness and model, and a person accepts every change.

Continual Learning

Review learning runs and candidate skill updates.

Add filterFilter by skill, run status, suggestion, ID, or date.

3 learning runs

RunStatusReplaysScoreCostCreated
Run lrn_2e91Quarterly portfolio review · 5 candidates · 1 suggested Done12/12Avg 96%$3.40Today
Run lrn_2d07Board pack extraction · 5 candidates · 1 suggested Done9/9Avg 91%$2.15Aug 21
Run lrn_2c88LP letter draft · 5 candidates · 0 suggested Replaying4/10Pending$0.90Aug 28
Page 1 of 1 PreviousNext
Quarterly portfolio review

Quarterly portfolio review

Candidate 3 · Today, 09:38

Status:
Done
Replays:
12/12 scored
Score:
96%
Cost:
$0.68

Run

Skill
Quarterly portfolio review
Run ID
lrn_2e91
Candidate ID
cand_7a2f
Completed
Today, 09:38

Candidate

Add the covenant chart from the workbook to each flagged company's slide.

Replay summary

Replay rows
12
Scored rows
12
Mean score
96%
Total cost
$0.68

Suggestion

PendingToday, 09:38

From the failed criterion on run_8f3a: slide 5 had no covenant chart. Every replayed deck now carries the chart beside each flagged company.

Quarterly portfolio review

Candidate 3 · Today, 09:38

Status:
Done
Replays:
12/12 scored
Score:
96%
Cost:
$0.68

Run

Skill
Quarterly portfolio review
Run ID
lrn_2e91
Candidate ID
cand_7a2f
Completed
Today, 09:38

Candidate

Add the covenant chart from the workbook to each flagged company's slide.

Replay summary

Replay rows
12
Scored rows
12
Mean score
96%
Total cost
$0.68

Suggestion

PendingToday, 09:38

From the failed criterion on run_8f3a: slide 5 had no covenant chart. Every replayed deck now carries the chart beside each flagged company.

Improve the procedures agents follow

Failed grades become candidate versions of a skill, tested on past tasks, and the best wait for a person to accept one.

Quarterly portfolio review

Candidate 3 · Today, 09:38

Status:
Done
Replays:
12/12 scored
Score:
96%
Cost:
$0.68

Run

Skill
Quarterly portfolio review
Run ID
lrn_2e91
Candidate ID
cand_7a2f
Completed
Today, 09:38

Candidate

Add the covenant chart from the workbook to each flagged company's slide.

Replay summary

Replay rows
12
Scored rows
12
Mean score
96%
Total cost
$0.68

Suggestion

PendingToday, 09:38

From the failed criterion on run_8f3a: slide 5 had no covenant chart. Every replayed deck now carries the chart beside each flagged company.

Tune the harness around each agent

Test changes to an agent's instructions, tools and settings against the current agent, on the same tasks and the same scorecard.

BenchmarksQuarterly review deckRun v4

Run v4 · Today 08:40

3 sources × 8 tasks · 24 executions

brun_4c1e · Agent judge · Completed in 2m 10s

CompletedRe-run
Pass rate88%21 of 24 passed
Progress24 / 2424 succeeded · 0 failed
Cost$14.97Complete run
Duration2m 10sStarted 3h ago
AtlasContext 1.5Quarterly portfolio review7 ok, 1 failedAtlasClaude Sonnet 5Quarterly portfolio review8 ok, 0 failedAtlasContext 1.5Quarterly portfolio review v26 ok, 2 failed

Results4 of 8 items · 3 sources

ItemAtlasContext 1.5AtlasClaude Sonnet 5AtlasContext 1.5
Task 1Build the Q3 portfolio review deck…PassPassFail
Task 2Build the Q2 portfolio review deck…PassPassPass
Task 3Build the Q1 portfolio review deck…FailPassFail
Task 4Build the Q4 2025 review deck…PassPassPass

Put the best model behind each agent

Benchmark models on your own tasks, compare pass rate and cost, and set the winner as the agent's default model.

Dashboards

Quarter to date

Edit Add widget
Pass ratefeedback · by week
60%80%100%97%
Spendllm_calls · $599 this quarter
$25$50$75

Grade every run against your checklist

Every run is recorded step by step and graded against the criteria your reviewers wrote. The grades decide what gets improved next.

Q3 portfolio review deck

Atlas · Maya Okafor

SummaryTimelineThreadFeedback
Duration37.5s
LLM Calls4
Sub-agents0
Tokens41.2k
Session overview37.5s active · 4 calls
  • 1 Read the updates doc and the workbook6.2s · 8.4k tok
  • 2 Draft slides from the template14.8s · 12.1k tok
  • 3 Place flagged companies first9.1s · 9.7k tok
  • 4 Render and post to the channel7.4s · 11.0k tok

Q3 portfolio review deck

Atlas · Maya Okafor

SummaryTimelineThreadFeedback
Duration37.5s
LLM Calls4
Sub-agents0
Tokens41.2k
Session overview37.5s active · 4 calls
  • 1 Read the updates doc and the workbook6.2s · 8.4k tok
  • 2 Draft slides from the template14.8s · 12.1k tok
  • 3 Place flagged companies first9.1s · 9.7k tok
  • 4 Render and post to the channel7.4s · 11.0k tok

Get a better skill from failed runs

A learning run reads the failed criteria, writes candidate versions of the skill, tests them on past tasks and keeps the best for your review.

Quarterly portfolio review

Candidate 3 · Today, 09:38

Status:
Done
Replays:
12/12 scored
Score:
96%
Cost:
$0.68

Run

Skill
Quarterly portfolio review
Run ID
lrn_2e91
Candidate ID
cand_7a2f
Completed
Today, 09:38

Candidate

Add the covenant chart from the workbook to each flagged company's slide.

Replay summary

Replay rows
12
Scored rows
12
Mean score
96%
Total cost
$0.68

Suggestion

PendingToday, 09:38

From the failed criterion on run_8f3a: slide 5 had no covenant chart. Every replayed deck now carries the chart beside each flagged company.

Quarterly portfolio review

Candidate 3 · Today, 09:38

Status:
Done
Replays:
12/12 scored
Score:
96%
Cost:
$0.68

Run

Skill
Quarterly portfolio review
Run ID
lrn_2e91
Candidate ID
cand_7a2f
Completed
Today, 09:38

Candidate

Add the covenant chart from the workbook to each flagged company's slide.

Replay summary

Replay rows
12
Scored rows
12
Mean score
96%
Total cost
$0.68

Suggestion

PendingToday, 09:38

From the failed criterion on run_8f3a: slide 5 had no covenant chart. Every replayed deck now carries the chart beside each flagged company.

Tune the harness around each agent

The harness is everything around the model: the agent's instructions, skills, connectors, permission rules and settings. Like a skill, it is text your team can read, change and test.

Q3 portfolio review deck Open in Drive Files Activity Share
You09:20

Build the Q3 review deck from the updates doc and the valuation bridge. Same structure as Q2, the three flagged companies up front.

Context09:20

Reading the Q3 updates and the bridge workbook, then drafting from the Q2 template.

Q3 portfolio review deck
Done in 3 min 48 s
  1. Read q3-portfolio-updates.docx, 14 companies
  2. Read q3-valuation-bridge.xlsx, 3 scenarios
  3. Draft slides from the Q2 template, 18 slides
  4. Place the flagged companies first, Northwind, Fenwick, Alder
  5. Render q3-portfolio-review.pptx, 4.1 MB
  6. Wrote q3-portfolio-review.pptx, 4.1 MB
Wrote q3-portfolio-review.pptx, 4.1 MB
q3-portfolio-review.pptxworkspaces/b/q3-portfolio-review-deckDeck · 4.1 MB
Context09:24

Deck is ready: 18 slides, the three variance companies on slides 3 to 5 with the covenant tests beside each, the full portfolio table in the appendix.

q3-portfolio-review.pptx4.1 MB · 18 slides
You09:31

Add the covenant headroom chart from the workbook to the Fenwick slide.

Context09:31

Added. Slide 5 now carries the interest-cover chart with the 2.0x test line; the appendix table is unchanged.

Share the deck with #q3-portfolio-review?
Posts q3-portfolio-review.pptx to the channel and notifies the four members.
ShareNot yet
Ask anything
CAuto
q3-portfolio-review.pptx
Deck

Q3 portfolio review deck

Sep 10, 2026 · Q3 portfolio review deck · Context

Deck is ready: 18 slides, the three variance companies on slides 3 to 5 with the covenant tests beside each, the full portfolio table in the appendix.

Added. Slide 5 now carries the interest-cover chart with the 2.0x test line; the appendix table is unchanged.

What was checked

Read q3-portfolio-updates.docx: 14 companies

Read q3-valuation-bridge.xlsx: 3 scenarios

Draft slides from the Q2 template: 18 slides

Q3 portfolio review deck Files Share
You09:20

Build the Q3 review deck from the updates doc and the valuation bridge. Same structure as Q2, the three flagged companies up front.

Context09:20

Reading the Q3 updates and the bridge workbook, then drafting from the Q2 template.

Q3 portfolio review deck
Done in 3 min 48 s
  1. Read q3-portfolio-updates.docx, 14 companies
  2. Read q3-valuation-bridge.xlsx, 3 scenarios
  3. Draft slides from the Q2 template, 18 slides
  4. Place the flagged companies first, Northwind, Fenwick, Alder
  5. Render q3-portfolio-review.pptx, 4.1 MB
  6. Wrote q3-portfolio-review.pptx, 4.1 MB
Wrote q3-portfolio-review.pptx, 4.1 MB
q3-portfolio-review.pptxworkspaces/b/q3-portfolio-review-deckDeck · 4.1 MB
Context09:24

Deck is ready: 18 slides, the three variance companies on slides 3 to 5 with the covenant tests beside each, the full portfolio table in the appendix.

q3-portfolio-review.pptx4.1 MB · 18 slides
You09:31

Add the covenant headroom chart from the workbook to the Fenwick slide.

Context09:31

Added. Slide 5 now carries the interest-cover chart with the 2.0x test line; the appendix table is unchanged.

Share the deck with #q3-portfolio-review?
Posts q3-portfolio-review.pptx to the channel and notifies the four members.
ShareNot yet
Ask anything
CAuto

Put the best model behind each agent

Run the same agent on different models against your own tasks, compare the grades and the cost, and switch the agent to the model that wins.

Quarterly review deck

Agent judge

New run Edit

Instructions

Compare the deck with the expected answer. Pass when the flagged companies, their covenants and headroom match; fail on any missing or invented figure.
Datasets Link dataset
Quarterly review promptsOrganization datasetDetach
Tasks8 Add task
IDPromptFilesExpected outputEdited
Task 1Build the Q3 portfolio review deck from the updates doc and the valuation bridge.2SetSep 12
Task 2Build the Q2 portfolio review deck from the updates doc and the valuation bridge.2SetSep 12
Task 3Build the Q1 portfolio review deck, with the two new investments marked.2SetSep 12
Details
Model
Configured on judge agent
Judge
Agent (sandboxed)
Datasets
1
Runs
3
Created
Sep 12, 2026
Updated
Today

Quarterly review deck

Agent judge

New run Edit

Instructions

Compare the deck with the expected answer. Pass when the flagged companies, their covenants and headroom match; fail on any missing or invented figure.
Datasets Link dataset
Quarterly review promptsOrganization datasetDetach
Tasks8 Add task
IDPromptFilesExpected outputEdited
Task 1Build the Q3 portfolio review deck from the updates doc and the valuation bridge.2SetSep 12
Task 2Build the Q2 portfolio review deck from the updates doc and the valuation bridge.2SetSep 12
Task 3Build the Q1 portfolio review deck, with the two new investments marked.2SetSep 12
Details
Model
Configured on judge agent
Judge
Agent (sandboxed)
Datasets
1
Runs
3
Created
Sep 12, 2026
Updated
Today

Nothing changes until a person accepts it

Grading, tests and learning runs happen inside your cluster. A person approves every change, and every earlier skill version can be restored.

Quarterly portfolio review
Quarterly portfolio reviewSavedv32/3 Improve RunSave
Quarterly portfolio review

Build the quarterly review deck from the updates doc and the valuation bridge, flagged companies first.

Scope
Brindle Capital
Tags
portfolio
Files
SKILL.md deck-template.pptx Add files
Connectors
+ Attach
Scorecards
Deck quality

Steps

  1. Read the quarter's updates doc and the valuation bridge workbook.
  2. Draft the slides from last quarter's template, one company per slide.
  3. Place every company more than 10% off plan first, on slides 3 to 5.
  4. Render the deck and post it to the review channel with the flagged companies named.

Written by

Atlas · from the Q3 review deck task

Quarterly portfolio review
Quarterly portfolio reviewSavedv32/3 Improve RunSave
Quarterly portfolio review

Build the quarterly review deck from the updates doc and the valuation bridge, flagged companies first.

Scope
Brindle Capital
Tags
portfolio
Files
SKILL.md deck-template.pptx Add files
Connectors
+ Attach
Scorecards
Deck quality

Steps

  1. Read the quarter's updates doc and the valuation bridge workbook.
  2. Draft the slides from last quarter's template, one company per slide.
  3. Place every company more than 10% off plan first, on slides 3 to 5.
  4. Render the deck and post it to the review channel with the flagged companies named.

Written by

Atlas · from the Q3 review deck task

Review suggested changes on any device

Review runs, grades and suggested versions on web, phone and desktop, and read runs and benchmarks through the CLI and the REST API.

Dashboards

Quarter to date

Edit Add widget
Pass ratefeedback · by week
60%80%100%97%
Spendllm_calls · $599 this quarter
$25$50$75

Dashboards

Quarter to date

Edit Add widget
Pass ratefeedback · by week
60%80%100%97%
Spendllm_calls · $599 this quarter
$25$50$75

Work across Context with shared permissions

Your runs, grades, skills and agents share one permission system, so your team can read every change Context RL suggests.

Home

Fri, Sep 11 72°
RecentPinnedTasks
  • q3-portfolio-review.pptx1h ago
  • q3-valuation-bridge.xlsx2h ago
  • q3-portfolio-updates.docx3h ago
  • lp-letter.pptx5h ago
Working on Q3 portfolio review deck

Home

Fri, Sep 11 72°
RecentPinnedTasks
  • q3-portfolio-review.pptx1h ago
  • q3-valuation-bridge.xlsx2h ago
  • q3-portfolio-updates.docx3h ago
  • lp-letter.pptx5h ago
Working on Q3 portfolio review deck

Deploy in your environment

Run grading, tests and learning runs inside your cluster, under the same deployment as your agents.

  • Run records stay in your cluster
  • Tests run in your cluster
  • Self-hosted through KOTS
  • SAML, OIDC and SCIM
How deployment works
  • Inside your clusterRun records, judges, tests and learning runs stay inside your cluster.
  • Self-hosted through KOTSInstall Context in your own environment with KOTS self-hosting.
  • Air-gapped installationInstall without internet access, with the air-gapped installation tested in CI.
  • Your identity providerSign people in with SAML or OIDC and provision them with SCIM.

Context runs in production at Qualcomm. Read the case study

  • Your own cloud
  • Your cluster
  • Air-gapped

Context RL is included in every Context workspace, including the Free plan.

  • Free$0Includes Context RL
  • Plus$20Per person per month
  • EnterpriseCustomEnterprise deployments run in your own cloud.Talk to us about a plan

Questions

What does Context RL improve?

Three things. The skills your agents follow, the harness around each agent (its instructions, skills, connectors, permission rules and settings) and the model each agent runs on. Every change is tested against your scorecards before a person accepts it.

What does it learn from?

It learns from your team's scorecard grades. A judge grades each run against your criteria, and the grades are collected for the skill version and agent that produced the run.

How is a change tested?

Candidate skill versions rerun past tasks through the task API and are graded by the same judge. Harness and model changes run as benchmark sources beside the current agent, on the same tasks and the same scorecard.

How do I pick a better model for an agent?

Run one benchmark with the same agent on several models, compare pass rate and cost for each, and set the best one as the agent's default model. Each run uses a fixed benchmark version, so later runs compare on the same tasks.

Can anything change without review?

No. Suggested skill versions wait for a person to accept them, and every accepted version can be restored. Harness and model changes take effect only when someone on your team saves them.

Can I take the data out?

Yes. Run records and grades stream as line-delimited JSON, and every read path is a GET endpoint under the Evals REST API with a scoped key.

Improve your agents' skills, harness and model with Context RL