Context EvalsProduction evals for agent workflows
Score every agent run
Your reviewers write scorecards, judges open each run's files, and scores feed back into the agent's skill.
Review is Thursday. Priya, can you get the portfolio pack started? Same structure as Q2, and Northwind and Alder go in the risk section this time.
@Scout pull the Q3 updates for all 14 portfolio companies from the data room and flag anything more than 10% off plan.
Three companies moved more than 10% against plan: Northwind Logistics (revenue down 14%, one lost contract), Alder Health (up 18%, the payer contract started early) and Fenwick Tools (EBITDA down 11%, freight). Sources are linked in the doc.
@Ledger rebuild the valuation bridge with the Q3 actuals and rerun the covenant tests for Northwind and Fenwick.
Northwind passes leverage at 3.9x against a 4.5x covenant. Fenwick's interest cover is 2.1x with the test at 2.0x: passing, but tight. One tab per company in the workbook.
@Atlas can we have the review deck from this, flagged companies up front?
Analytics
Review evaluation records, feedback, and run outcomes.
| Task | Workspace | User | Agent | Skills | Feedback | Created |
|---|---|---|---|---|---|---|
| Q3 portfolio review deckrun_8f3a | Brindle Capital | Maya Okafor | Atlas | Quarterly portfolio review | 2.0 (1) | Today 09:24 |
| Covenant headroom appletrun_8f41 | Brindle Capital | Marco Vitale | Ledger | None | None | Today 09:33 |
| Q3 valuation bridgerun_8f2c | Brindle Capital | Priya Natarajan | Ledger | Covenant test rerun | 5.0 (2) | Today 09:17 |
| Q3 portfolio updatesrun_8f11 | Brindle Capital | Maya Okafor | Scout | Board pack extraction | 4.5 (2) | Today 09:10 |
| LP letter · Q2run_7c90 | Brindle Capital | Maya Okafor | Atlas | LP letter draft | 4.0 (1) | Aug 28 |
- 1The criterion that failed
- 2The Improver's proposal
- 3Replayed and accepted
Score every run in production
Every completed agent run is scored against its applicable scorecard, so you can measure quality on the work itself.
Dashboards
Quarter to date
Evaluate the files agents produce
An agent judge opens each run's files and trace in its own sandbox and returns a verdict and reason for each criterion.
Q3 portfolio review deck
Atlas · Brindle Capital
Maya Okafor · Today 09:24
- Workspace:
- Brindle Capital
- Avg Rating:
- 2.0
- Feedback:
- 1
- Messages:
- 6
- Runs:
- 1
- Attempts:
- 1
Prompt
# Where this task came fromYou were dispatched by Atlas, a Together agent, from the public channel #q3-portfolio-review, at the request of Maya Okafor.
Feedback
- Flagged companies lead the deckYes
- Covenant test beside each flagged companyNoPriya Natarajan · Fenwick's slide has no covenant chart.
- Appendix table completeYes
- Overall2 / 5
Additional FeedbackFenwick's slide has no covenant chart. The workbook's Fenwick tab holds the test.
Improve on the same model
At Qualcomm, agents went from a 23% to a 98% task pass rate in four months, with no fine-tuning and no model change.
Quarterly portfolio review
Candidate 3 · Today, 09:38
- Status:
- Done
- Replays:
- 12/12 scored
- Score:
- 96%
- Cost:
- $0.68
Run
Candidate
Add the covenant chart from the workbook to each flagged company's slide.Replay summary
Suggestion
From the failed criterion on run_8f3a: slide 5 had no covenant chart. Every replayed deck now carries the chart beside each flagged company.
Score every run against your criteria
Write your criteria once as a scorecard to score every run in its scope, with the verdict displayed beside each run.
Q3 portfolio review deck
Atlas · Brindle Capital
Maya Okafor · Today 09:24
- Workspace:
- Brindle Capital
- Avg Rating:
- 2.0
- Feedback:
- 1
- Messages:
- 6
- Runs:
- 1
- Attempts:
- 1
Prompt
# Where this task came fromYou were dispatched by Atlas, a Together agent, from the public channel #q3-portfolio-review, at the request of Maya Okafor.
Feedback
- Flagged companies lead the deckYes
- Covenant test beside each flagged companyNoPriya Natarajan · Fenwick's slide has no covenant chart.
- Appendix table completeYes
- Overall2 / 5
Additional FeedbackFenwick's slide has no covenant chart. The workbook's Fenwick tab holds the test.
Q3 portfolio review deck
Atlas · Brindle Capital
Maya Okafor · Today 09:24
- Workspace:
- Brindle Capital
- Avg Rating:
- 2.0
- Feedback:
- 1
- Messages:
- 6
- Runs:
- 1
- Attempts:
- 1
Prompt
# Where this task came fromYou were dispatched by Atlas, a Together agent, from the public channel #q3-portfolio-review, at the request of Maya Okafor.
Feedback
- Flagged companies lead the deckYes
- Covenant test beside each flagged companyNoPriya Natarajan · Fenwick's slide has no covenant chart.
- Appendix table completeYes
- Overall2 / 5
Additional FeedbackFenwick's slide has no covenant chart. The workbook's Fenwick tab holds the test.
Benchmark agents, skills and models
Run your prompt dataset against any agent, skill and model, using a judge that reads the output or opens the deliverables.
Quarterly review deck
Agent judge
Instructions
Compare the deck with the expected answer. Pass when the flagged companies, their covenants and headroom match; fail on any missing or invented figure.| ID | Prompt | Files | Expected output | Edited |
|---|---|---|---|---|
| Task 1 | Build the Q3 portfolio review deck from the updates doc and the valuation bridge. | 2 | Set | Sep 12 |
| Task 2 | Build the Q2 portfolio review deck from the updates doc and the valuation bridge. | 2 | Set | Sep 12 |
| Task 3 | Build the Q1 portfolio review deck, with the two new investments marked. | 2 | Set | Sep 12 |
Quarterly review deck
Agent judge
Instructions
Compare the deck with the expected answer. Pass when the flagged companies, their covenants and headroom match; fail on any missing or invented figure.| ID | Prompt | Files | Expected output | Edited |
|---|---|---|---|---|
| Task 1 | Build the Q3 portfolio review deck from the updates doc and the valuation bridge. | 2 | Set | Sep 12 |
| Task 2 | Build the Q2 portfolio review deck from the updates doc and the valuation bridge. | 2 | Set | Sep 12 |
| Task 3 | Build the Q1 portfolio review deck, with the two new investments marked. | 2 | Set | Sep 12 |
Track quality and improve skills
Read each run's trace and track pass rate and cost per agent, with failed criteria informing proposals for the skill's next version.
Q3 portfolio review deck
Atlas · Maya Okafor
- 1 Read the updates doc and the workbook
- 2 Draft slides from the template
- 3 Place flagged companies first
- 4 Render and post to the channel
Q3 portfolio review deck
Atlas · Maya Okafor
- 1 Read the updates doc and the workbook
- 2 Draft slides from the template
- 3 Place flagged companies first
- 4 Render and post to the channel
Control access by default
Choose which organization roles can open Evals, with the judge in its own sandbox and read access limited to the run it scores.
BenchmarksQuarterly review deckRun v4
Run v4 · Today 08:40
3 sources × 8 tasks · 24 executions
brun_4c1e · Agent judge · Completed in 2m 10s
Results
| Item | AtlasContext 1.5 | AtlasClaude Sonnet 5 | AtlasContext 1.5 |
|---|---|---|---|
| Task 1Build the Q3 portfolio review deck… | Pass | Pass | Fail |
| Task 2Build the Q2 portfolio review deck… | Pass | Pass | Pass |
| Task 3Build the Q1 portfolio review deck… | Fail | Pass | Fail |
| Task 4Build the Q4 2025 review deck… | Pass | Pass | Pass |
BenchmarksQuarterly review deckRun v4
Run v4 · Today 08:40
3 sources × 8 tasks · 24 executions
brun_4c1e · Agent judge · Completed in 2m 10s
Results
| Item | AtlasContext 1.5 | AtlasClaude Sonnet 5 | AtlasContext 1.5 |
|---|---|---|---|
| Task 1Build the Q3 portfolio review deck… | Pass | Pass | Fail |
| Task 2Build the Q2 portfolio review deck… | Pass | Pass | Pass |
| Task 3Build the Q1 portfolio review deck… | Fail | Pass | Fail |
| Task 4Build the Q4 2025 review deck… | Pass | Pass | Pass |
Read evals across your devices
Open dashboards, analytics, benchmarks and traces on web, phone and desktop, search traces and read benchmarks through the CLI, and query that data through the REST API.
Analytics
Review evaluation records, feedback, and run outcomes.
| Task | Workspace | User | Agent | Skills | Feedback | Created |
|---|---|---|---|---|---|---|
| Q3 portfolio review deckrun_8f3a | Brindle Capital | Maya Okafor | Atlas | Quarterly portfolio review | 2.0 (1) | Today 09:24 |
| Covenant headroom appletrun_8f41 | Brindle Capital | Marco Vitale | Ledger | None | None | Today 09:33 |
| Q3 valuation bridgerun_8f2c | Brindle Capital | Priya Natarajan | Ledger | Covenant test rerun | 5.0 (2) | Today 09:17 |
| Q3 portfolio updatesrun_8f11 | Brindle Capital | Maya Okafor | Scout | Board pack extraction | 4.5 (2) | Today 09:10 |
| LP letter · Q2run_7c90 | Brindle Capital | Maya Okafor | Atlas | LP letter draft | 4.0 (1) | Aug 28 |
Q3 portfolio review deck
Atlas · Brindle Capital
Maya Okafor · Today 09:24
- Workspace:
- Brindle Capital
- Avg Rating:
- 2.0
- Feedback:
- 1
- Messages:
- 6
- Runs:
- 1
- Attempts:
- 1
Prompt
# Where this task came fromYou were dispatched by Atlas, a Together agent, from the public channel #q3-portfolio-review, at the request of Maya Okafor.
Feedback
- Flagged companies lead the deckYes
- Covenant test beside each flagged companyNoPriya Natarajan · Fenwick's slide has no covenant chart.
- Appendix table completeYes
- Overall2 / 5
Additional FeedbackFenwick's slide has no covenant chart. The workbook's Fenwick tab holds the test.
Q3 portfolio review deck
Atlas · Brindle Capital
Maya Okafor · Today 09:24
- Workspace:
- Brindle Capital
- Avg Rating:
- 2.0
- Feedback:
- 1
- Messages:
- 6
- Runs:
- 1
- Attempts:
- 1
Prompt
# Where this task came fromYou were dispatched by Atlas, a Together agent, from the public channel #q3-portfolio-review, at the request of Maya Okafor.
Feedback
- Flagged companies lead the deckYes
- Covenant test beside each flagged companyNoPriya Natarajan · Fenwick's slide has no covenant chart.
- Appendix table completeYes
- Overall2 / 5
Additional FeedbackFenwick's slide has no covenant chart. The workbook's Fenwick tab holds the test.
Work across the Context platform
Workspace is where your people and agents work, Engine runs it, Unify connects it to your tools, and Evals makes it better.
- WorkspaceA task started in chat or a channel is scored like any other, with feedback collected where the work happened.
- EngineBenchmark the models Engine runs side by side, with pass rate and cost recorded for each.
- UnifyRuns that use your connected apps and internal services are scored against the same scorecards as every other run.
- SkillsFailed criteria feed a learning run that writes the skill's next version for your review, with its replay score beside it.
Home
- q3-portfolio-review.pptx1h ago
- q3-valuation-bridge.xlsx2h ago
- q3-portfolio-updates.docx3h ago
- lp-letter.pptx5h ago
Home
- q3-portfolio-review.pptx1h ago
- q3-valuation-bridge.xlsx2h ago
- q3-portfolio-updates.docx3h ago
- lp-letter.pptx5h ago
Deploy in your environment
Run judges, traces, recordings and learning runs inside your cluster, under the same deployment as your agents.
- Inside your clusterJudges, traces, recordings and learning runs run inside your cluster.
- Self-hosted through KOTSInstall Context in your own environment with KOTS self-hosting.
- Air-gapped installationInstall without internet access, with the air-gapped installation tested in CI.
- Your identity providerSign people in with SAML or OIDC and provision them with SCIM.
At Qualcomm, the same model went from 23% to 98% task pass rate in four months. Read the case study
- Your own cloud
- Your cluster
- Air-gapped
Context Evals is included in every Context workspace, including the Free plan.
- Free$0Includes Context Evals
- Plus$20Per person per month
- EnterpriseCustomEnterprise deployments run in your own cloud.Talk to us about a plan
Questions
What is a scorecard?
A scorecard groups criteria into a reusable evaluation. Each criterion includes an output type (pass or fail, a score, or text), a required flag and instructions for the judge. A scorecard applies to an organization, a workspace or a person, and every change creates an immutable version.
What does the agent judge see?
The judge sees the run's deliverables and trace, copied into its own sandbox, and uses that evidence to score each criterion and write a verdict file with a reason per criterion. A malformed verdict triggers one more prompt, then an explicit failure if it remains malformed. A score is never fabricated.
Can I benchmark different models?
Yes, a benchmark run accepts up to twenty sources, each an agent with its skills and a model from that agent's configuration. It records pass rate and cost per source. Each run uses a fixed, immutable benchmark version, so results stay comparable.
How does continual learning change a skill?
Scorecard feedback is buffered per skill version. Once a batch collects, a learning run proposes five candidates, replays each against the buffered runs, scores them and keeps the best as a proposed next version, within cost and elapsed-time limits. A person reviews the proposed version before it is saved.
Who can open Evals?
Evals access is set by organization role. Owners and admins have access by default, and an admin can add members. Every eval read is authorized against the caller's organization.
Can I read eval data from outside the app?
Yes, every read path is a GET endpoint under the Evals REST API. Access requires a task API key with the evals read capability. Analytics records also stream as line-delimited JSON for warehouse syncs.



