Skip to main content

Test metadata

A promptfoo test already has a free-form metadata object. Hajer reserves one key inside it, hajer, and reads promptfoo's own testCaseId. Nothing else about the test changes.

promptfooconfig.yaml
defaultTest:
metadata:
hajer:
workflowId: wf_support # inherited by every test

tests:
- description: Pending refund is reported as pending
metadata:
testCaseId: eval_refund_pending # promptfoo's stable id for this test
hajer:
obligationIds: [obl_refund_status_disclosed]
componentIds: [cmp_refund_agent]
sourceTraceIds: [3f2a9c0d1e7b4c55a8e2f61b0c9d4e7a]
generatedBy: support-team
provenance:
ticket: SUP-1182
vars:
message: Has my refund gone through?
customer_id: cust_123
assert:
- type: trajectory:tool-used
value: get_refund_status

metadata.hajer keys​

KeyTypeRequiredMeaning
workflowIdstringyesThe workflow the test targets. The same id your code passes to hajer.workflow(...).
obligationIdslist of stringsnoThe obligations the test covers. Each must be declared in hajer.yaml when one is found.
componentIdslist of stringsnoComponents the test expects to exercise. The trace records which actually ran.
sourceTraceIdslist of stringsnoProduction traces this test was derived from.
generatedBystringnoWho or what wrote the test.
provenanceobjectnoAny further detail about where the test came from.

Limits on every id (workflowId and each list entry):

  • must be a string; a number or boolean is refused,
  • leading and trailing whitespace is trimmed,
  • must not be empty, and at most 128 characters after trimming,
  • ids within one list must be unique.

Keys are case-sensitive. workflowID is an unknown key, not a spelling of workflowId.

metadata.testCaseId​

testCaseId is promptfoo's own key for a stable test id. Hajer uses it to match a test's result to its trace and to line up the same test across runs.

Set it on every test that has metadata.hajer. Without it the test still runs, with a warning, and correlation falls back to the test's position in the file. That position changes when you reorder or insert tests, so history on the platform stops lining up.

hajer eval: warning [W_NO_TEST_CASE_ID] test #2 "An unknown account is told so, not guessed at" (workflowId=wf_support): no metadata.testCaseId; correlation falls back to the test's position, which moves when tests are reordered; add metadata.testCaseId

Warnings are also copied into the run's payload, so they appear with the run on the platform.

Which tests are correlated​

The test hasResult
no metadata.hajer, and no defaultTest.metadata.hajerAn ordinary promptfoo test. It runs unchanged and is reported with correlation: none.
a valid metadata.hajer (its own, inherited, or both)Correlated with the workflow: correlation: platform.
a valid metadata.hajer but no testCaseIdCorrelated, with the W_NO_TEST_CASE_ID warning.
a malformed metadata.hajerThe run stops before any provider is called. Exit code 2.

A hajer: key with no value (YAML null) counts as absent.

Inheritance from defaultTest​

promptfoo merges defaultTest.metadata into each test shallowly, so a test's own hajer object would replace the inherited one entirely. hajer eval merges the hajer object one level deeper first: each key from defaultTest.metadata.hajer is kept unless the test sets the same key.

defaultTest:
metadata:
hajer:
workflowId: wf_support

tests:
- metadata:
testCaseId: eval_a
hajer:
obligationIds: [obl_refund_status_disclosed]
# effective: { workflowId: wf_support, obligationIds: [obl_refund_status_disclosed] }
- metadata:
testCaseId: eval_b
hajer:
workflowId: wf_billing
# effective: { workflowId: wf_billing }
- metadata:
testCaseId: eval_c
# effective: { workflowId: wf_support }

The merge is per key, not per list entry. A test that sets obligationIds replaces the inherited list; it does not append to it.

caution

If defaultTest.metadata.hajer exists, every test inherits a hajer object. When that inherited object has no workflowId, every test that does not set one fails validation with metadata.hajer.workflowId is required. Put workflowId in defaultTest or on each test.

Errors​

Validation runs in a promptfoo beforeAll hook, after the suite is loaded and before the first provider call, so a mistake costs no model calls. Every error names the test by position and by its testCaseId (or description), and starts with a code you can search for.

CodeCauseEffect
E_HAJER_INVALIDAn unknown key, a missing workflowId, a wrong type, an empty or over-long id, or a repeated id in a list.Run stops, exit 2.
E_HAJER_NOT_OBJECTmetadata.hajer is a string, number, boolean or list instead of an object.Run stops, exit 2.
E_OBLIGATION_UNDECLAREDAn entry in obligationIds is not declared in the hajer.yaml in force.Run stops, exit 2.
W_NO_TEST_CASE_IDA correlated test has no metadata.testCaseId.Warning only; the test runs.

Example output for a misspelled key:

hajer eval: error [E_HAJER_INVALID] test #0 "t1": metadata.hajer.workflowId is required
hajer eval: error [E_HAJER_INVALID] test #0 "t1": metadata.hajer.workflowID is not a known key; the keys are workflowId, obligationIds, componentIds, sourceTraceIds, generatedBy, provenance

Filtering by metadata​

hajer eval --workflow ID and --obligation ID run only the tests whose effective metadata.hajer matches. See the CLI reference. promptfoo's own --filter-metadata reads top-level metadata keys only and cannot see inside hajer.