Skip to main content

hajer eval command

hajer eval [-c PATH ...] [--suite ID] [--workflow ID] [--obligation ID ...]
[--upload | --no-upload] [--payload-out PATH]
[--install-only] [--engine-help] [any promptfoo eval flag]

hajer eval is a subcommand of the hajer console script. python -m hajer eval is the same program.

Flags​

FlagMeaning
-c PATH, --config PATHRun this promptfoo configuration instead of the manifest's suites. Repeatable: all named files run together as one engine run. Relative paths resolve against the working directory.
--suite IDRun only the suite hajer.yaml declares under this id. Cannot be combined with -c.
--workflow IDRun only tests whose effective metadata.hajer.workflowId equals ID.
--obligation IDRun only tests whose metadata.hajer.obligationIds contains ID. Repeatable; a test matching any of them is kept.
--uploadUpload each run's payload to Hajer. Needs HAJER_API_KEY and HAJER_TEAM_ID. See Uploading runs.
--no-uploadDo not upload. This is the default; the flag exists so a script can say so explicitly. It wins over --upload.
--payload-out PATHAlso write the payload here. A file path for one suite; a directory, receiving <suite id>.json per suite, when several suites run.
--install-onlyInstall the pinned engine into the cache, print its status as JSON, and exit. See The engine.
--engine-helpPrint promptfoo eval --help for the pinned engine, which lists every flag you can pass through.
-h, --helpPrint hajer eval's own help.

Passing flags to promptfoo​

Any flag hajer eval does not recognise is passed to promptfoo eval unchanged, for example:

hajer eval --no-cache -j 4
hajer eval --suite support --filter-pattern "refund" --repeat 3

hajer eval adds its own -c for an overlay configuration and -o for the results file, so do not pass -o yourself; use --payload-out for the payload.

Suite selection​

Invocationhajer.yaml foundWhat runs
hajer evalyesEvery declared suite, in order, each as its own engine run, payload and upload. Each runs from the manifest's directory.
hajer eval --suite IDyesThat one suite. An undeclared id exits 2 and lists the declared ones.
hajer evalnopromptfoo's default configuration in the working directory: the first of promptfooconfig.yaml, .yml, .json, .js, .mjs, .cjs, .ts. None found exits 2.
hajer eval --suite IDnoExits 2.
hajer eval -c a.yaml [-c b.yaml]eitherThe named files, as one run, from the working directory. The manifest, if found, still declares which obligation ids tests may name.

For how the manifest is found, see Discovery.

Filtering tests​

--workflow and --obligation narrow the suite before the first provider call:

  • Both given: a test must match both.
  • Several --obligation: a test must cover at least one.
  • A test without metadata.hajer never matches an active filter.
  • Filters apply to the effective metadata, after defaultTest inheritance.
  • An --obligation the manifest does not declare is refused before the engine starts (exit 2).
  • If no test matches, the run stops with exit 2 and says how many tests the suite had.
hajer eval --obligation obl_refund_status_disclosed
hajer eval --workflow wf_support --obligation obl_refund_status_disclosed

The filter is recorded in the payload under filters.

Output​

Where payloads are written​

Every run gets its own directory under the cache:

$HAJER_CACHE_DIR/runs/<run id>/
├── overlay.json the configuration hajer eval adds (tracing, the beforeAll hook)
├── promptfoo/ promptfoo's config directory for this run
├── hook-report.json what the metadata check found
├── results.json promptfoo's own results document
└── payload.json the Hajer payload

HAJER_CACHE_DIR defaults to $XDG_CACHE_HOME/hajer when XDG_CACHE_HOME is set, otherwise ~/.cache/hajer. The newest 20 run directories are kept (HAJER_EVAL_RUNS_KEEP); older ones are removed when a run starts.

A run id looks like evalrun_ followed by 32 hex characters. It is created before the engine starts.

With --payload-out, the payload is written there as well, and the summary line points at that copy.

The summary line​

Each suite ends with one line on stdout:

hajer eval (suite support): passed: 4/4 passed, 0 failed, 0 errored; run evalrun_c4bc72e7657048a0998423495bff21be; payload /home/you/.cache/hajer/runs/evalrun_c4bc72e7657048a0998423495bff21be/payload.json; upload not requested

The parts are: the suite (or plain hajer eval for a -c run), the run status (passed, failed, errored or aborted), the test counts, the run id, the payload path, and the upload outcome. The upload outcome is one of not requested, uploaded, skipped (INERT), or failed (REASON). When the upload did not succeed, a line on stderr repeats where the payload is kept.

promptfoo's own output (the results table and its summary) is printed above it, unchanged. Metadata warnings and errors are printed after promptfoo exits.

Exit codes​

CodeMeaningStops a multi-suite run
0Every test passed.no
100At least one test failed or errored. This is promptfoo's own code, passed through.no; the remaining suites still run
2Usage: no configuration found, an invalid hajer.yaml, invalid metadata.hajer, an undeclared obligation, --suite combined with -c, an unknown suite id, or no test matched a filter.yes
3Node.js or npm is missing, or Node.js is older than 22.22.0.checked before any suite runs
4The engine could not be installed or failed its smoke test.checked before any suite runs
5The engine exited 0 but wrote no readable results document.yes

Any other non-zero code promptfoo returns is passed through unchanged.

For a manifest run, the command exits with the highest code among the suites that ran. A code 2 or 5 ends the run at that suite; suites after it do not run.

info

Uploading never affects the exit code. A skipped upload (no credentials) and a failed upload (refused, unreachable, too large) are reported on the summary line and stderr, and the command still exits with the engine's verdict.

Examples​

hajer eval # every suite in hajer.yaml
hajer eval --suite support # one suite
hajer eval -c evals/experimental.yaml # a file not in the manifest
hajer eval --payload-out out/ # copies of each payload in out/<suite id>.json
hajer eval --upload # run and report to Hajer
hajer eval --install-only # warm the engine cache
hajer eval --engine-help # promptfoo eval's own flags