Test metadata
A promptfoo test already has a free-form metadata object. Hajer reserves one key inside it, hajer, and reads
promptfoo's own testCaseId. Nothing else about the test changes.
defaultTest:
metadata:
hajer:
workflowId: wf_support # inherited by every test
tests:
- description: Pending refund is reported as pending
metadata:
testCaseId: eval_refund_pending # promptfoo's stable id for this test
hajer:
obligationIds: [obl_refund_status_disclosed]
componentIds: [cmp_refund_agent]
sourceTraceIds: [3f2a9c0d1e7b4c55a8e2f61b0c9d4e7a]
generatedBy: support-team
provenance:
ticket: SUP-1182
vars:
message: Has my refund gone through?
customer_id: cust_123
assert:
- type: trajectory:tool-used
value: get_refund_status
metadata.hajer keys
| Key | Type | Required | Meaning |
|---|---|---|---|
workflowId | string | yes | The workflow the test targets. The same id your code passes to hajer.workflow(...). |
obligationIds | list of strings | no | The obligations the test covers. Each must be declared in hajer.yaml when one is found. |
componentIds | list of strings | no | Components the test expects to exercise. The trace records which actually ran. |
sourceTraceIds | list of strings | no | Production traces this test was derived from. |
generatedBy | string | no | Who or what wrote the test. |
provenance | object | no | Any further detail about where the test came from. |
Limits on every id (workflowId and each list entry):
- must be a string; a number or boolean is refused,
- leading and trailing whitespace is trimmed,
- must not be empty, and at most 128 characters after trimming,
- ids within one list must be unique.
Keys are case-sensitive. workflowID is an unknown key, not a spelling of workflowId.
metadata.testCaseId
testCaseId is promptfoo's own key for a stable test id. Hajer uses it to match a test's result to its trace and
to line up the same test across runs.
Set it on every test that has metadata.hajer. Without it the test still runs, with a warning, and correlation
falls back to the test's position in the file. That position changes when you reorder or insert tests, so
history on the platform stops lining up.
hajer eval: warning [W_NO_TEST_CASE_ID] test #2 "An unknown account is told so, not guessed at" (workflowId=wf_support): no metadata.testCaseId; correlation falls back to the test's position, which moves when tests are reordered; add metadata.testCaseId
Warnings are also copied into the run's payload, so they appear with the run on the platform.
Which tests are correlated
| The test has | Result |
|---|---|
no metadata.hajer, and no defaultTest.metadata.hajer | An ordinary promptfoo test. It runs unchanged and is reported with correlation: none. |
a valid metadata.hajer (its own, inherited, or both) | Correlated with the workflow: correlation: platform. |
a valid metadata.hajer but no testCaseId | Correlated, with the W_NO_TEST_CASE_ID warning. |
a malformed metadata.hajer | The run stops before any provider is called. Exit code 2. |
A hajer: key with no value (YAML null) counts as absent.
Inheritance from defaultTest
promptfoo merges defaultTest.metadata into each test shallowly, so a test's own hajer object would replace the
inherited one entirely. hajer eval merges the hajer object one level deeper first: each key from
defaultTest.metadata.hajer is kept unless the test sets the same key.
defaultTest:
metadata:
hajer:
workflowId: wf_support
tests:
- metadata:
testCaseId: eval_a
hajer:
obligationIds: [obl_refund_status_disclosed]
# effective: { workflowId: wf_support, obligationIds: [obl_refund_status_disclosed] }
- metadata:
testCaseId: eval_b
hajer:
workflowId: wf_billing
# effective: { workflowId: wf_billing }
- metadata:
testCaseId: eval_c
# effective: { workflowId: wf_support }
The merge is per key, not per list entry. A test that sets obligationIds replaces the inherited list; it does not
append to it.
If defaultTest.metadata.hajer exists, every test inherits a hajer object. When that inherited object has no
workflowId, every test that does not set one fails validation with metadata.hajer.workflowId is required.
Put workflowId in defaultTest or on each test.
Errors
Validation runs in a promptfoo beforeAll hook, after the suite is loaded and before the first provider call, so
a mistake costs no model calls. Every error names the test by position and by its testCaseId (or description),
and starts with a code you can search for.
| Code | Cause | Effect |
|---|---|---|
E_HAJER_INVALID | An unknown key, a missing workflowId, a wrong type, an empty or over-long id, or a repeated id in a list. | Run stops, exit 2. |
E_HAJER_NOT_OBJECT | metadata.hajer is a string, number, boolean or list instead of an object. | Run stops, exit 2. |
E_OBLIGATION_UNDECLARED | An entry in obligationIds is not declared in the hajer.yaml in force. | Run stops, exit 2. |
W_NO_TEST_CASE_ID | A correlated test has no metadata.testCaseId. | Warning only; the test runs. |
Example output for a misspelled key:
hajer eval: error [E_HAJER_INVALID] test #0 "t1": metadata.hajer.workflowId is required
hajer eval: error [E_HAJER_INVALID] test #0 "t1": metadata.hajer.workflowID is not a known key; the keys are workflowId, obligationIds, componentIds, sourceTraceIds, generatedBy, provenance
Filtering by metadata
hajer eval --workflow ID and --obligation ID run only the tests whose effective metadata.hajer matches. See
the CLI reference. promptfoo's own --filter-metadata reads top-level metadata keys
only and cannot see inside hajer.