Reinventing.AI
Insights / Harnesses

Compare the work, not the demo.

Use this protocol to test a harness on your own workflow. This is a published test method and worksheet, not a report of benchmark results.

Define one representative task

Use the same input, allowed actions, deliverable, and acceptance criteria for each harness. A useful starting task is preparing a sourced draft from a fixed set of permitted documents. Keep external delivery disabled during the comparison.

Record the harness and model versions, operating system, tool configuration, permission mode, instruction revision, and resource limits. If model choice differs, describe the result as a comparison of complete configurations rather than isolating the effect of the harness.

Use a small, disclosed test set

Include normal input, missing input, conflicting evidence, a tool failure, and an input that attempts to change the instructions. Repeat each case enough times to see variability and report the actual sample size. Preserve failed attempts alongside successful ones.

The downloadable CSV contains blank result fields. Fill them only after running the task. Use sanitized fixtures and share only artifacts you have permission to publish.

Measure operational outcomes

  • Acceptance: did the output satisfy the task contract?
  • Boundary behavior: did the run attempt or perform an unauthorized action?
  • Interventions: what did a person have to fix or approve?
  • Recovery: could the task resume from a verified checkpoint?
  • Cost: what were model, tool, and allocated runtime costs?
  • Time: how long did execution and human review each take?

Report these dimensions separately. An aggregate score can hide a serious boundary failure behind a fast completion time. If a vendor does not expose a cost or execution trace, record it as unavailable rather than estimating it without a stated method.

Publish limitations with the result

A small test of one workflow cannot establish which harness is universally best. Explain the task, configuration, dates, sample size, missing data, and manual interventions. Link the permitted input fixtures and output artifacts so another operator can understand the result.

Re-run relevant cases after changing the model, harness, tools, or permissions. Keep earlier results dated rather than silently presenting them as current. For additional test design guidance, see Anthropic’s evaluation guidance.

Keep the comparison connected to the job

Use the same acceptance criteria after launch, when real inputs become less predictable.

Build your workflow evaluation