TL;DR
Pick a research task, read the whole plan before anything runs, approve a compute ceiling, and stop at the decision point the plan puts in front of the expensive half.
Use it / Skip it
Workbench is the plan-first entry to the platform. A research task expands into an explicit, numbered plan; every step names the tool it will call and how many metered jobs it will submit. Nothing runs until you approve, and approving is also how you set the ceiling. One step in the plan is a decision point: the campaign stops there and waits for you to look at the cheap output before the expensive jobs are committed.
Use when
You want a complete design-and-rank chain rather than a single tool run, and you want to see the cost and the tool list before spending anything.
Don't use for
Do not use it as a shortcut around a target that has no measured cohort. Workbench will still run, but the ranking is a triage order, not a predicted affinity — read the applicability-domain note on the plan.
Inputs
- Research task
- One of the starter cards, or a goal described in the agent.
- Target
- Resolved to a structure and epitope; the plan states where the epitope came from.
- Compute ceiling
- The number of metered jobs the campaign may spend. Set at approval time.
Outputs
- Plan
- Numbered steps, each with its tool id and metered job count.
- Campaign
- A persisted row with status, spend against ceiling, and one job id per step.
- Shortlist
- Ranked candidates with every contributing number and its provenance attached.
- Audit
- A deterministic review that reads the recorded measurements and reports what cannot be defended.
Walkthrough
1. Pick a starter task
Pick a starter task. The card shows the plan and the total metered job count before you commit.
2. Press Create this plan for approval
Press Create this plan for approval. The campaign is created but nothing has run.
3. Set the ceiling and approve
Set the ceiling and approve. Approval and budget are the same action, so a plan cannot run without a limit.
4. Run to next decision point
Run to next decision point. The campaign executes the cheap steps and stops at the review step.
5. Read the cheap output, then narrow
Read the cheap output, then narrow. Steps you drop are never submitted, so the saving is real.
6. Audit this run when the chain finishes, and read the shortlist with its provenance column
Audit this run when the chain finishes, and read the shortlist with its provenance column.
Under the hood
- Campaign, steps, approval and spend are database rows; the ceiling is enforced by a database trigger, not by UI code.
- Each step submits through the shared tool contract, so the same plan runs against whichever executor a tool is bound to.
- Step status is derived from job status, so a page reload never invents progress.
- The reviewer is deterministic: it reads recorded measurements, so the same run always produces the same findings.
Worked example
Pitfalls
- Approving without lowering the ceiling gives away the decision point's whole benefit.
- A step marked skipped is not a failure — open it and read the reason before rerunning.
- Ranking is on the interface and developability rubric. It is not a predicted K_D unless the target has a calibrated cohort.