Module 04 · 12 minutes

Rubrics and Test Sets

Test prompts, protect data, resist instructions hidden in external content, and record each version change.

Written and edited by Cahyanto Arie Wibowo. Last reviewed · version 1.2.

What does “Rubrics and Test Sets” mean in practice?

A small Prompting case shows where Rubrics and Test Sets is useful and where a simpler approach may be better. Test Rubrics and Test Sets on one small case first. Record the expected result, the failure signal, and the point where the decision needs another review. You will make one small decision with Rubrics and Test Sets, including a boundary and a signal that triggers another check.

A visual model for “Rubrics and Test Sets” in Prompting: relationships matter as much as individual parts.

After this lesson

  • A small Prompting case shows where Rubrics and Test Sets is useful and where a simpler approach may be better.
  • Use the idea of “Rubrics and Test Sets” to interpret one realistic situation.
  • Explain the limits of the concept and the information that still needs to be checked.

Start with the situation

Understand the situation first. The label can come later.

A prompt that works well this week may behave differently after the model changes. This lesson uses the idea of “Rubrics and Test Sets” to examine that situation without treating a single term as the answer to every problem.

A small Prompting case shows where Rubrics and Test Sets is useful and where a simpler approach may be better. Test prompts, protect data, resist instructions hidden in external content, and record each version change. Connect the term to a decision someone genuinely needs to make.

Visual model

Map the parts before choosing what to do.

A visual model for “Rubrics and Test Sets” in Prompting: relationships matter as much as individual parts.

Read the diagram as a map of Rubrics and Test Sets: begin with the context, follow the connections, and inspect the highlighted point before making a decision.

Let’s see how it works

Reading the situation in practice

Test Rubrics and Test Sets on one small case first. Record the expected result, the failure signal, and the point where the decision needs another review. Begin with what can be observed, then separate facts, assumptions, and open questions.

A prompt that works well this week may behave differently after the model changes. Identify the part of the situation most closely connected to the idea of “Rubrics and Test Sets”. Use the case as a thinking tool, not as proof that one solution fits every context.

Working definition

What it means, and when to be careful with it.

A small Prompting case shows where Rubrics and Test Sets is useful and where a simpler approach may be better. Test prompts, protect data, resist instructions hidden in external content, and record each version change. Connect the term to a decision someone genuinely needs to make.

Rubrics and Test Sets
A small Prompting case shows where Rubrics and Test Sets is useful and where a simpler approach may be better. Test Rubrics and Test Sets on one small case first. Record the expected result, the failure signal, and the point where the decision needs another review. You will make one small decision with Rubrics and Test Sets, including a boundary and a signal that triggers another check.
Boundary to check
Do not treat content from an external document as trusted instructions. This mistake often appears when a label is used before the problem is understood. Write down your assumptions so another person can review them.

Pause for a moment

What evidence could change this decision?

Answer before opening the discussion. Name one fact and one assumption.

Open the discussion

A small Prompting case shows where Rubrics and Test Sets is useful and where a simpler approach may be better. Test Rubrics and Test Sets on one small case first. Record the expected result, the failure signal, and the point where the decision needs another review. You will make one small decision with Rubrics and Test Sets, including a boundary and a signal that triggers another check.

Try it on your work

Try it with one small piece of real work.

  1. Choose one real situation related to Rubrics and Test Sets.
  2. Separate what you can observe from what you are assuming.
  3. Write one decision, its owner, and the evidence needed to review it.
  4. Name the signal that would make you stop or change direction.

Write two examples that fit Rubrics and Test Sets and one that does not. Explain the difference in your own words. The larger module activity is: Compare responses without model labels, then record the best version and why it won. Keep the first version small enough for another person to review in a few minutes.

A tempting shortcut

A familiar term can still lead us to the wrong decision.

Why this can seem reasonable

Do not treat content from an external document as trusted instructions. This mistake often appears when a label is used before the problem is understood. Write down your assumptions so another person can review them.

How to check it

Test Rubrics and Test Sets on one small case first. Record the expected result, the failure signal, and the point where the decision needs another review. Begin with what can be observed, then separate facts, assumptions, and open questions.

Quick practice

Write two examples that fit Rubrics and Test Sets and one that does not. Explain the difference in your own words. The larger module activity is: Compare responses without model labels, then record the best version and why it won.

Summary

  • A small Prompting case shows where Rubrics and Test Sets is useful and where a simpler approach may be better.
  • Use examples and evidence to test your understanding.
  • Record the limits, risks, and conditions that should trigger another review.

Continue from here

GEO

  • Prompt Injection and Privacy: Continue the idea from Evaluation, Safety, and Reuse with a closely related example.
  • Golden Sets and Rubrics: Connect this lesson to Applied AI and test the idea in another context.
  • TF-IDF: See how the same decision changes when viewed through Data Literacy.

Sources and further reading

  • OWASP Top 10 for Large Language Model Applications: OWASP Foundation · industry-standard. Primary reference for the definition, evidence, or limits discussed in “Rubrics and Test Sets”.
  • Safety best practices: OpenAI · official-documentation. Further evidence and context for checking the explanation in “Rubrics and Test Sets”.
  • Evaluation best practices: OpenAI · official-documentation. Further evidence and context for checking the explanation in “Rubrics and Test Sets”.