Context EvalsProduction evals for agent workflows

Score every agent run

Your reviewers write scorecards, judges open each run's files, and scores feed back into the agent's skill.

q3-portfolio-reviewThursday's review: updates, bridge, deck, covenant watch7
Messages FilesDetails
Today
Maya Okafor09:02

Review is Thursday. Priya, can you get the portfolio pack started? Same structure as Q2, and Northwind and Alder go in the risk section this time.

Priya Natarajan09:04

@Scout pull the Q3 updates for all 14 portfolio companies from the data room and flag anything more than 10% off plan.

Scout09:04
Q3 portfolio updates
Done in 6 min 12 s
Wrote q3-portfolio-updates.docx, 184 KB
q3-portfolio-updates.docxworkspaces/b/q3-portfolio-reviewDoc · 184 KB

Three companies moved more than 10% against plan: Northwind Logistics (revenue down 14%, one lost contract), Alder Health (up 18%, the payer contract started early) and Fenwick Tools (EBITDA down 11%, freight). Sources are linked in the doc.

2
Marco Vitale09:12

@Ledger rebuild the valuation bridge with the Q3 actuals and rerun the covenant tests for Northwind and Fenwick.

Ledger09:13
Q3 valuation bridge
Done in 4 min 40 s
Wrote q3-valuation-bridge.xlsx, 612 KB
q3-valuation-bridge.xlsxworkspaces/b/q3-portfolio-reviewSheet · 612 KB

Northwind passes leverage at 3.9x against a 4.5x covenant. Fenwick's interest cover is 2.1x with the test at 2.0x: passing, but tight. One tab per company in the workbook.

Maya Okafor09:20

@Atlas can we have the review deck from this, flagged companies up front?

Atlas09:21
Q3 portfolio review deck
Done in 3 min 48 s
Wrote q3-portfolio-review.pptx, 4.1 MB
q3-portfolio-review.pptxworkspaces/b/q3-portfolio-reviewDeck · 4.1 MB
3
Send a message in #q3-portfolio-review

Analytics

Review evaluation records, feedback, and run outcomes.

Export CSV
View Date | 2026-06-30 → 2026-09-19 Add filterSave view
Total Tasks323Top-level tasks
Total Records341Rows
Total Feedback217Feedback
Avg Rating4.3 / 5Average
Data 341Graphs
TaskWorkspaceUserAgentSkillsFeedbackCreated
Q3 portfolio review deckrun_8f3aBrindle CapitalMaya OkaforAtlasQuarterly portfolio review2.0 (1)Today 09:24
Covenant headroom appletrun_8f41Brindle CapitalMarco VitaleLedgerNoneNoneToday 09:33
Q3 valuation bridgerun_8f2cBrindle CapitalPriya NatarajanLedgerCovenant test rerun5.0 (2)Today 09:17
Q3 portfolio updatesrun_8f11Brindle CapitalMaya OkaforScoutBoard pack extraction4.5 (2)Today 09:10
LP letter · Q2run_7c90Brindle CapitalMaya OkaforAtlasLP letter draft4.0 (1)Aug 28
  1. 1The criterion that failed
  2. 2The Improver's proposal
  3. 3Replayed and accepted

Score every run in production

Every completed agent run is scored against its applicable scorecard, so you can measure quality on the work itself.

Dashboards

Quarter to date

Edit Add widget
Pass ratefeedback · by week
60%80%100%97%
Spendllm_calls · $599 this quarter
$25$50$75

Evaluate the files agents produce

An agent judge opens each run's files and trace in its own sandbox and returns a verdict and reason for each criterion.

Q3 portfolio review deck

Atlas · Brindle Capital

Maya Okafor · Today 09:24

Workspace:
Brindle Capital
Avg Rating:
2.0
Feedback:
1
Messages:
6
Runs:
1
Attempts:
1
View Chat
Skills used:Quarterly portfolio review

Prompt

# Where this task came fromYou were dispatched by Atlas, a Together agent, from the public channel #q3-portfolio-review, at the request of Maya Okafor.

FeedbackExport Feedback CSV

Deck qualityProject FeedbackToday, 09:41
  • Flagged companies lead the deckYes
  • Covenant test beside each flagged companyNoPriya Natarajan · Fenwick's slide has no covenant chart.
  • Appendix table completeYes
  • Overall2 / 5

Additional FeedbackFenwick's slide has no covenant chart. The workbook's Fenwick tab holds the test.

Improve on the same model

At Qualcomm, agents went from a 23% to a 98% task pass rate in four months, with no fine-tuning and no model change.

Quarterly portfolio review

Candidate 3 · Today, 09:38

Status:
Done
Replays:
12/12 scored
Score:
96%
Cost:
$0.68

Run

Skill
Quarterly portfolio review
Run ID
lrn_2e91
Candidate ID
cand_7a2f
Completed
Today, 09:38

Candidate

Add the covenant chart from the workbook to each flagged company's slide.

Replay summary

Replay rows
12
Scored rows
12
Mean score
96%
Total cost
$0.68

Suggestion

PendingToday, 09:38

From the failed criterion on run_8f3a: slide 5 had no covenant chart. Every replayed deck now carries the chart beside each flagged company.

Score every run against your criteria

Write your criteria once as a scorecard to score every run in its scope, with the verdict displayed beside each run.

Q3 portfolio review deck

Atlas · Brindle Capital

Maya Okafor · Today 09:24

Workspace:
Brindle Capital
Avg Rating:
2.0
Feedback:
1
Messages:
6
Runs:
1
Attempts:
1
View Chat
Skills used:Quarterly portfolio review

Prompt

# Where this task came fromYou were dispatched by Atlas, a Together agent, from the public channel #q3-portfolio-review, at the request of Maya Okafor.

FeedbackExport Feedback CSV

Deck qualityProject FeedbackToday, 09:41
  • Flagged companies lead the deckYes
  • Covenant test beside each flagged companyNoPriya Natarajan · Fenwick's slide has no covenant chart.
  • Appendix table completeYes
  • Overall2 / 5

Additional FeedbackFenwick's slide has no covenant chart. The workbook's Fenwick tab holds the test.

Q3 portfolio review deck

Atlas · Brindle Capital

Maya Okafor · Today 09:24

Workspace:
Brindle Capital
Avg Rating:
2.0
Feedback:
1
Messages:
6
Runs:
1
Attempts:
1
View Chat
Skills used:Quarterly portfolio review

Prompt

# Where this task came fromYou were dispatched by Atlas, a Together agent, from the public channel #q3-portfolio-review, at the request of Maya Okafor.

FeedbackExport Feedback CSV

Deck qualityProject FeedbackToday, 09:41
  • Flagged companies lead the deckYes
  • Covenant test beside each flagged companyNoPriya Natarajan · Fenwick's slide has no covenant chart.
  • Appendix table completeYes
  • Overall2 / 5

Additional FeedbackFenwick's slide has no covenant chart. The workbook's Fenwick tab holds the test.

Benchmark agents, skills and models

Run your prompt dataset against any agent, skill and model, using a judge that reads the output or opens the deliverables.

Quarterly review deck

Agent judge

New run Edit

Instructions

Compare the deck with the expected answer. Pass when the flagged companies, their covenants and headroom match; fail on any missing or invented figure.
Datasets Link dataset
Quarterly review promptsOrganization datasetDetach
Tasks8 Add task
IDPromptFilesExpected outputEdited
Task 1Build the Q3 portfolio review deck from the updates doc and the valuation bridge.2SetSep 12
Task 2Build the Q2 portfolio review deck from the updates doc and the valuation bridge.2SetSep 12
Task 3Build the Q1 portfolio review deck, with the two new investments marked.2SetSep 12
Details
Model
Configured on judge agent
Judge
Agent (sandboxed)
Datasets
1
Runs
3
Created
Sep 12, 2026
Updated
Today

Quarterly review deck

Agent judge

New run Edit

Instructions

Compare the deck with the expected answer. Pass when the flagged companies, their covenants and headroom match; fail on any missing or invented figure.
Datasets Link dataset
Quarterly review promptsOrganization datasetDetach
Tasks8 Add task
IDPromptFilesExpected outputEdited
Task 1Build the Q3 portfolio review deck from the updates doc and the valuation bridge.2SetSep 12
Task 2Build the Q2 portfolio review deck from the updates doc and the valuation bridge.2SetSep 12
Task 3Build the Q1 portfolio review deck, with the two new investments marked.2SetSep 12
Details
Model
Configured on judge agent
Judge
Agent (sandboxed)
Datasets
1
Runs
3
Created
Sep 12, 2026
Updated
Today

Track quality and improve skills

Read each run's trace and track pass rate and cost per agent, with failed criteria informing proposals for the skill's next version.

Q3 portfolio review deck

Atlas · Maya Okafor

SummaryTimelineThreadFeedback
Duration37.5s
LLM Calls4
Sub-agents0
Tokens41.2k
Session overview37.5s active · 4 calls
  • 1 Read the updates doc and the workbook6.2s · 8.4k tok
  • 2 Draft slides from the template14.8s · 12.1k tok
  • 3 Place flagged companies first9.1s · 9.7k tok
  • 4 Render and post to the channel7.4s · 11.0k tok

Q3 portfolio review deck

Atlas · Maya Okafor

SummaryTimelineThreadFeedback
Duration37.5s
LLM Calls4
Sub-agents0
Tokens41.2k
Session overview37.5s active · 4 calls
  • 1 Read the updates doc and the workbook6.2s · 8.4k tok
  • 2 Draft slides from the template14.8s · 12.1k tok
  • 3 Place flagged companies first9.1s · 9.7k tok
  • 4 Render and post to the channel7.4s · 11.0k tok

Control access by default

Choose which organization roles can open Evals, with the judge in its own sandbox and read access limited to the run it scores.

BenchmarksQuarterly review deckRun v4

Run v4 · Today 08:40

3 sources × 8 tasks · 24 executions

brun_4c1e · Agent judge · Completed in 2m 10s

CompletedRe-run
Pass rate88%21 of 24 passed
Progress24 / 2424 succeeded · 0 failed
Cost$14.97Complete run
Duration2m 10sStarted 3h ago
AtlasContext 1.5Quarterly portfolio review7 ok, 1 failedAtlasClaude Sonnet 5Quarterly portfolio review8 ok, 0 failedAtlasContext 1.5Quarterly portfolio review v26 ok, 2 failed

Results4 of 8 items · 3 sources

ItemAtlasContext 1.5AtlasClaude Sonnet 5AtlasContext 1.5
Task 1Build the Q3 portfolio review deck…PassPassFail
Task 2Build the Q2 portfolio review deck…PassPassPass
Task 3Build the Q1 portfolio review deck…FailPassFail
Task 4Build the Q4 2025 review deck…PassPassPass

BenchmarksQuarterly review deckRun v4

Run v4 · Today 08:40

3 sources × 8 tasks · 24 executions

brun_4c1e · Agent judge · Completed in 2m 10s

CompletedRe-run
Pass rate88%21 of 24 passed
Progress24 / 2424 succeeded · 0 failed
Cost$14.97Complete run
Duration2m 10sStarted 3h ago
AtlasContext 1.5Quarterly portfolio review7 ok, 1 failedAtlasClaude Sonnet 5Quarterly portfolio review8 ok, 0 failedAtlasContext 1.5Quarterly portfolio review v26 ok, 2 failed

Results4 of 8 items · 3 sources

ItemAtlasContext 1.5AtlasClaude Sonnet 5AtlasContext 1.5
Task 1Build the Q3 portfolio review deck…PassPassFail
Task 2Build the Q2 portfolio review deck…PassPassPass
Task 3Build the Q1 portfolio review deck…FailPassFail
Task 4Build the Q4 2025 review deck…PassPassPass

Read evals across your devices

Open dashboards, analytics, benchmarks and traces on web, phone and desktop, search traces and read benchmarks through the CLI, and query that data through the REST API.

Analytics

Review evaluation records, feedback, and run outcomes.

Export CSV
View Date | 2026-06-30 → 2026-09-19 Add filterSave view
Total Tasks323Top-level tasks
Total Records341Rows
Total Feedback217Feedback
Avg Rating4.3 / 5Average
Data 341Graphs
TaskWorkspaceUserAgentSkillsFeedbackCreated
Q3 portfolio review deckrun_8f3aBrindle CapitalMaya OkaforAtlasQuarterly portfolio review2.0 (1)Today 09:24
Covenant headroom appletrun_8f41Brindle CapitalMarco VitaleLedgerNoneNoneToday 09:33
Q3 valuation bridgerun_8f2cBrindle CapitalPriya NatarajanLedgerCovenant test rerun5.0 (2)Today 09:17
Q3 portfolio updatesrun_8f11Brindle CapitalMaya OkaforScoutBoard pack extraction4.5 (2)Today 09:10
LP letter · Q2run_7c90Brindle CapitalMaya OkaforAtlasLP letter draft4.0 (1)Aug 28
341 recordsPreviousNext
Q3 portfolio review deck

Q3 portfolio review deck

Atlas · Brindle Capital

Maya Okafor · Today 09:24

Workspace:
Brindle Capital
Avg Rating:
2.0
Feedback:
1
Messages:
6
Runs:
1
Attempts:
1
View Chat
Skills used:Quarterly portfolio review

Prompt

# Where this task came fromYou were dispatched by Atlas, a Together agent, from the public channel #q3-portfolio-review, at the request of Maya Okafor.

FeedbackExport Feedback CSV

Deck qualityProject FeedbackToday, 09:41
  • Flagged companies lead the deckYes
  • Covenant test beside each flagged companyNoPriya Natarajan · Fenwick's slide has no covenant chart.
  • Appendix table completeYes
  • Overall2 / 5

Additional FeedbackFenwick's slide has no covenant chart. The workbook's Fenwick tab holds the test.

Q3 portfolio review deck

Atlas · Brindle Capital

Maya Okafor · Today 09:24

Workspace:
Brindle Capital
Avg Rating:
2.0
Feedback:
1
Messages:
6
Runs:
1
Attempts:
1
View Chat
Skills used:Quarterly portfolio review

Prompt

# Where this task came fromYou were dispatched by Atlas, a Together agent, from the public channel #q3-portfolio-review, at the request of Maya Okafor.

FeedbackExport Feedback CSV

Deck qualityProject FeedbackToday, 09:41
  • Flagged companies lead the deckYes
  • Covenant test beside each flagged companyNoPriya Natarajan · Fenwick's slide has no covenant chart.
  • Appendix table completeYes
  • Overall2 / 5

Additional FeedbackFenwick's slide has no covenant chart. The workbook's Fenwick tab holds the test.

Work across the Context platform

Workspace is where your people and agents work, Engine runs it, Unify connects it to your tools, and Evals makes it better.

Home

Fri, Sep 11 72°
RecentPinnedTasks
  • q3-portfolio-review.pptx1h ago
  • q3-valuation-bridge.xlsx2h ago
  • q3-portfolio-updates.docx3h ago
  • lp-letter.pptx5h ago
Working on Q3 portfolio review deck

Home

Fri, Sep 11 72°
RecentPinnedTasks
  • q3-portfolio-review.pptx1h ago
  • q3-valuation-bridge.xlsx2h ago
  • q3-portfolio-updates.docx3h ago
  • lp-letter.pptx5h ago
Working on Q3 portfolio review deck

Deploy in your environment

Run judges, traces, recordings and learning runs inside your cluster, under the same deployment as your agents.

  • Judges run in your cluster
  • Traces and recordings in your cluster
  • Self-hosted through KOTS
  • SAML, OIDC and SCIM
How deployment works
  • Inside your clusterJudges, traces, recordings and learning runs run inside your cluster.
  • Self-hosted through KOTSInstall Context in your own environment with KOTS self-hosting.
  • Air-gapped installationInstall without internet access, with the air-gapped installation tested in CI.
  • Your identity providerSign people in with SAML or OIDC and provision them with SCIM.

At Qualcomm, the same model went from 23% to 98% task pass rate in four months. Read the case study

  • Your own cloud
  • Your cluster
  • Air-gapped

Context Evals is included in every Context workspace, including the Free plan.

  • Free$0Includes Context Evals
  • Plus$20Per person per month
  • EnterpriseCustomEnterprise deployments run in your own cloud.Talk to us about a plan

Questions

What is a scorecard?

A scorecard groups criteria into a reusable evaluation. Each criterion includes an output type (pass or fail, a score, or text), a required flag and instructions for the judge. A scorecard applies to an organization, a workspace or a person, and every change creates an immutable version.

What does the agent judge see?

The judge sees the run's deliverables and trace, copied into its own sandbox, and uses that evidence to score each criterion and write a verdict file with a reason per criterion. A malformed verdict triggers one more prompt, then an explicit failure if it remains malformed. A score is never fabricated.

Can I benchmark different models?

Yes, a benchmark run accepts up to twenty sources, each an agent with its skills and a model from that agent's configuration. It records pass rate and cost per source. Each run uses a fixed, immutable benchmark version, so results stay comparable.

How does continual learning change a skill?

Scorecard feedback is buffered per skill version. Once a batch collects, a learning run proposes five candidates, replays each against the buffered runs, scores them and keeps the best as a proposed next version, within cost and elapsed-time limits. A person reviews the proposed version before it is saved.

Who can open Evals?

Evals access is set by organization role. Owners and admins have access by default, and an admin can add members. Every eval read is authorized against the caller's organization.

Can I read eval data from outside the app?

Yes, every read path is a GET endpoint under the Evals REST API. Access requires a task API key with the evals read capability. Analytics records also stream as line-delimited JSON for warehouse syncs.

Score every agent run with Context Evals