The idea
A test set should reflect the inputs and failures the workflow actually encounters. Include ordinary cases, rare but costly failures, missing information and unsupported requests. Record where samples came from and whether you may use them. A convenient set of easy prompts can give a misleading quality score.
Worked example
A support-drafting evaluation contains clear questions, vague questions and refund requests the assistant cannot authorize. If all examples are simple delivery questions, the measured success rate says little about how it handles unsupported actions or missing order details.
Try it
Write eight fictional inputs for one narrow drafting task. Label the situation each represents and include two costly failure cases. Explain which real-world input types remain missing and why your sample cannot support a claim of universal reliability.
