Context RLBetter skills, harness and model from graded runs
Improve your agents from every graded run
Your team grades agent runs with scorecards. Context uses the grades to improve each agent's skills, harness and model, and a person accepts every change.
Continual Learning
Review learning runs and candidate skill updates.
3 learning runs
| Run | Status | Replays | Score | Cost | Created |
|---|---|---|---|---|---|
| Run lrn_2e91Quarterly portfolio review · 5 candidates · 1 suggested | Done | 12/12 | Avg 96% | $3.40 | Today |
| Run lrn_2d07Board pack extraction · 5 candidates · 1 suggested | Done | 9/9 | Avg 91% | $2.15 | Aug 21 |
| Run lrn_2c88LP letter draft · 5 candidates · 0 suggested | Replaying | 4/10 | Pending | $0.90 | Aug 28 |
Quarterly portfolio review
Candidate 3 · Today, 09:38
- Status:
- Done
- Replays:
- 12/12 scored
- Score:
- 96%
- Cost:
- $0.68
Run
Candidate
Add the covenant chart from the workbook to each flagged company's slide.Replay summary
Suggestion
From the failed criterion on run_8f3a: slide 5 had no covenant chart. Every replayed deck now carries the chart beside each flagged company.
Quarterly portfolio review
Candidate 3 · Today, 09:38
- Status:
- Done
- Replays:
- 12/12 scored
- Score:
- 96%
- Cost:
- $0.68
Run
Candidate
Add the covenant chart from the workbook to each flagged company's slide.Replay summary
Suggestion
From the failed criterion on run_8f3a: slide 5 had no covenant chart. Every replayed deck now carries the chart beside each flagged company.
Improve the procedures agents follow
Failed grades become candidate versions of a skill, tested on past tasks, and the best wait for a person to accept one.
Quarterly portfolio review
Candidate 3 · Today, 09:38
- Status:
- Done
- Replays:
- 12/12 scored
- Score:
- 96%
- Cost:
- $0.68
Run
Candidate
Add the covenant chart from the workbook to each flagged company's slide.Replay summary
Suggestion
From the failed criterion on run_8f3a: slide 5 had no covenant chart. Every replayed deck now carries the chart beside each flagged company.
Tune the harness around each agent
Test changes to an agent's instructions, tools and settings against the current agent, on the same tasks and the same scorecard.
BenchmarksQuarterly review deckRun v4
Run v4 · Today 08:40
3 sources × 8 tasks · 24 executions
brun_4c1e · Agent judge · Completed in 2m 10s
Results
| Item | AtlasContext 1.5 | AtlasClaude Sonnet 5 | AtlasContext 1.5 |
|---|---|---|---|
| Task 1Build the Q3 portfolio review deck… | Pass | Pass | Fail |
| Task 2Build the Q2 portfolio review deck… | Pass | Pass | Pass |
| Task 3Build the Q1 portfolio review deck… | Fail | Pass | Fail |
| Task 4Build the Q4 2025 review deck… | Pass | Pass | Pass |
Put the best model behind each agent
Benchmark models on your own tasks, compare pass rate and cost, and set the winner as the agent's default model.
Dashboards
Quarter to date
Grade every run against your checklist
Every run is recorded step by step and graded against the criteria your reviewers wrote. The grades decide what gets improved next.
Q3 portfolio review deck
Atlas · Maya Okafor
- 1 Read the updates doc and the workbook
- 2 Draft slides from the template
- 3 Place flagged companies first
- 4 Render and post to the channel
Q3 portfolio review deck
Atlas · Maya Okafor
- 1 Read the updates doc and the workbook
- 2 Draft slides from the template
- 3 Place flagged companies first
- 4 Render and post to the channel
Get a better skill from failed runs
A learning run reads the failed criteria, writes candidate versions of the skill, tests them on past tasks and keeps the best for your review.
Quarterly portfolio review
Candidate 3 · Today, 09:38
- Status:
- Done
- Replays:
- 12/12 scored
- Score:
- 96%
- Cost:
- $0.68
Run
Candidate
Add the covenant chart from the workbook to each flagged company's slide.Replay summary
Suggestion
From the failed criterion on run_8f3a: slide 5 had no covenant chart. Every replayed deck now carries the chart beside each flagged company.
Quarterly portfolio review
Candidate 3 · Today, 09:38
- Status:
- Done
- Replays:
- 12/12 scored
- Score:
- 96%
- Cost:
- $0.68
Run
Candidate
Add the covenant chart from the workbook to each flagged company's slide.Replay summary
Suggestion
From the failed criterion on run_8f3a: slide 5 had no covenant chart. Every replayed deck now carries the chart beside each flagged company.
Tune the harness around each agent
The harness is everything around the model: the agent's instructions, skills, connectors, permission rules and settings. Like a skill, it is text your team can read, change and test.
Build the Q3 review deck from the updates doc and the valuation bridge. Same structure as Q2, the three flagged companies up front.
Reading the Q3 updates and the bridge workbook, then drafting from the Q2 template.
- Read q3-portfolio-updates.docx, 14 companies
- Read q3-valuation-bridge.xlsx, 3 scenarios
- Draft slides from the Q2 template, 18 slides
- Place the flagged companies first, Northwind, Fenwick, Alder
- Render q3-portfolio-review.pptx, 4.1 MB
- Wrote q3-portfolio-review.pptx, 4.1 MB
Deck is ready: 18 slides, the three variance companies on slides 3 to 5 with the covenant tests beside each, the full portfolio table in the appendix.
q3-portfolio-review.pptx4.1 MB · 18 slidesAdd the covenant headroom chart from the workbook to the Fenwick slide.
Added. Slide 5 now carries the interest-cover chart with the 2.0x test line; the appendix table is unchanged.
Q3 portfolio review deck
Deck is ready: 18 slides, the three variance companies on slides 3 to 5 with the covenant tests beside each, the full portfolio table in the appendix.
Added. Slide 5 now carries the interest-cover chart with the 2.0x test line; the appendix table is unchanged.
What was checked
Read q3-portfolio-updates.docx: 14 companies
Read q3-valuation-bridge.xlsx: 3 scenarios
Draft slides from the Q2 template: 18 slides
How access works
Atlas works in the channels it is a member of and reads the workspace Drive. Runs that need more ask first.
Skills
Quarterly portfolio review · Board pack extraction · LP letter draft
Build the Q3 review deck from the updates doc and the valuation bridge. Same structure as Q2, the three flagged companies up front.
Reading the Q3 updates and the bridge workbook, then drafting from the Q2 template.
- Read q3-portfolio-updates.docx, 14 companies
- Read q3-valuation-bridge.xlsx, 3 scenarios
- Draft slides from the Q2 template, 18 slides
- Place the flagged companies first, Northwind, Fenwick, Alder
- Render q3-portfolio-review.pptx, 4.1 MB
- Wrote q3-portfolio-review.pptx, 4.1 MB
Deck is ready: 18 slides, the three variance companies on slides 3 to 5 with the covenant tests beside each, the full portfolio table in the appendix.
q3-portfolio-review.pptx4.1 MB · 18 slidesAdd the covenant headroom chart from the workbook to the Fenwick slide.
Added. Slide 5 now carries the interest-cover chart with the 2.0x test line; the appendix table is unchanged.
How access works
Atlas works in the channels it is a member of and reads the workspace Drive. Runs that need more ask first.
Skills
Quarterly portfolio review · Board pack extraction · LP letter draft
Put the best model behind each agent
Run the same agent on different models against your own tasks, compare the grades and the cost, and switch the agent to the model that wins.
Quarterly review deck
Agent judge
Instructions
Compare the deck with the expected answer. Pass when the flagged companies, their covenants and headroom match; fail on any missing or invented figure.| ID | Prompt | Files | Expected output | Edited |
|---|---|---|---|---|
| Task 1 | Build the Q3 portfolio review deck from the updates doc and the valuation bridge. | 2 | Set | Sep 12 |
| Task 2 | Build the Q2 portfolio review deck from the updates doc and the valuation bridge. | 2 | Set | Sep 12 |
| Task 3 | Build the Q1 portfolio review deck, with the two new investments marked. | 2 | Set | Sep 12 |
Quarterly review deck
Agent judge
Instructions
Compare the deck with the expected answer. Pass when the flagged companies, their covenants and headroom match; fail on any missing or invented figure.| ID | Prompt | Files | Expected output | Edited |
|---|---|---|---|---|
| Task 1 | Build the Q3 portfolio review deck from the updates doc and the valuation bridge. | 2 | Set | Sep 12 |
| Task 2 | Build the Q2 portfolio review deck from the updates doc and the valuation bridge. | 2 | Set | Sep 12 |
| Task 3 | Build the Q1 portfolio review deck, with the two new investments marked. | 2 | Set | Sep 12 |
Nothing changes until a person accepts it
Grading, tests and learning runs happen inside your cluster. A person approves every change, and every earlier skill version can be restored.
Version history
- Version 3Current
- Version 2Restore
- Version 1Restore
Build the quarterly review deck from the updates doc and the valuation bridge, flagged companies first.
- Scope
- Brindle Capital
- Tags
- portfolio
- Files
- SKILL.md deck-template.pptx Add files
- Connectors
- + Attach
- Scorecards
- Deck quality
Steps
- Read the quarter's updates doc and the valuation bridge workbook.
- Draft the slides from last quarter's template, one company per slide.
- Place every company more than 10% off plan first, on slides 3 to 5.
- Render the deck and post it to the review channel with the flagged companies named.
Written by
Atlas · from the Q3 review deck task
Version history
- Version 3Current
- Version 2Restore
- Version 1Restore
Build the quarterly review deck from the updates doc and the valuation bridge, flagged companies first.
- Scope
- Brindle Capital
- Tags
- portfolio
- Files
- SKILL.md deck-template.pptx Add files
- Connectors
- + Attach
- Scorecards
- Deck quality
Steps
- Read the quarter's updates doc and the valuation bridge workbook.
- Draft the slides from last quarter's template, one company per slide.
- Place every company more than 10% off plan first, on slides 3 to 5.
- Render the deck and post it to the review channel with the flagged companies named.
Written by
Atlas · from the Q3 review deck task
Review suggested changes on any device
Review runs, grades and suggested versions on web, phone and desktop, and read runs and benchmarks through the CLI and the REST API.
Dashboards
Quarter to date
Dashboards
Quarter to date
Work across Context with shared permissions
Your runs, grades, skills and agents share one permission system, so your team can read every change Context RL suggests.
- EvalsScorecards, judges and benchmarks produce the grades that decide what Context RL improves.
- SkillsEach accepted fix is a new skill version, kept with its history and assigned to the agents that run it.
- WikiThe standards and procedures your agents follow sit beside your team's pages, editable by the people who own them.
- DriveThe judge opens the files a run wrote to Drive in its own sandbox, without reading the rest of your storage.
Home
- q3-portfolio-review.pptx1h ago
- q3-valuation-bridge.xlsx2h ago
- q3-portfolio-updates.docx3h ago
- lp-letter.pptx5h ago
Home
- q3-portfolio-review.pptx1h ago
- q3-valuation-bridge.xlsx2h ago
- q3-portfolio-updates.docx3h ago
- lp-letter.pptx5h ago
Deploy in your environment
Run grading, tests and learning runs inside your cluster, under the same deployment as your agents.
- Inside your clusterRun records, judges, tests and learning runs stay inside your cluster.
- Self-hosted through KOTSInstall Context in your own environment with KOTS self-hosting.
- Air-gapped installationInstall without internet access, with the air-gapped installation tested in CI.
- Your identity providerSign people in with SAML or OIDC and provision them with SCIM.
Context runs in production at Qualcomm. Read the case study
- Your own cloud
- Your cluster
- Air-gapped
Context RL is included in every Context workspace, including the Free plan.
- Free$0Includes Context RL
- Plus$20Per person per month
- EnterpriseCustomEnterprise deployments run in your own cloud.Talk to us about a plan
Questions
What does Context RL improve?
Three things. The skills your agents follow, the harness around each agent (its instructions, skills, connectors, permission rules and settings) and the model each agent runs on. Every change is tested against your scorecards before a person accepts it.
What does it learn from?
It learns from your team's scorecard grades. A judge grades each run against your criteria, and the grades are collected for the skill version and agent that produced the run.
How is a change tested?
Candidate skill versions rerun past tasks through the task API and are graded by the same judge. Harness and model changes run as benchmark sources beside the current agent, on the same tasks and the same scorecard.
How do I pick a better model for an agent?
Run one benchmark with the same agent on several models, compare pass rate and cost for each, and set the best one as the agent's default model. Each run uses a fixed benchmark version, so later runs compare on the same tasks.
Can anything change without review?
No. Suggested skill versions wait for a person to accept them, and every accepted version can be restored. Harness and model changes take effect only when someone on your team saves them.
Can I take the data out?
Yes. Run records and grades stream as line-delimited JSON, and every read path is a GET endpoint under the Evals REST API with a scoped key.



