Mental model
A capability-fit experiment is a sequence of falsifiable gates: offline task evidence, shadow workflow, bounded pilot, and monitored production decision.
Learning outcomes
- Explain the mechanism and ownership boundaries behind Capability-Fit Experimentation.
- Compare the main design alternatives and their operational trade-offs.
- Diagnose common failures and select evidence for a production decision.
Theory
Define baseline, cohort, success and harm metrics, stop conditions, representative data, human review, cost, latency, and version fingerprints. Separate model ability from integration and adoption effects.
Trade-offs
Offline tests are controlled but miss workflow behavior. Live pilots reveal adoption and operations but require stronger containment and causal caution.
Failure modes and misconceptions
No baseline; cherry-picked tasks; changing model and workflow together; missing stop conditions; measuring output preference instead of task outcome; and scaling before failure analysis.
Decision scenario
An extraction model scores well offline but reviewers still redo most records. Design the next experiment to locate capability, interface, or workflow mismatch.
Why should capability, workflow value, and operational fit be evaluated as separate hypotheses?
Primary sources
openai-evalsnist-ai-rmf