The engine
hajer eval runs promptfoo, unmodified, at one exact version. Hajer
adds no JavaScript. Everything it adds is Python, attached through promptfoo's public extension points: a
beforeAll hook, an extra configuration file, promptfoo's local trace receiver, the traceparent promptfoo hands
to providers, and the results file.
The pinned version
hajer 0.2.0 | |
|---|---|
| Engine | promptfoo 0.123.1, from the npm registry (MIT) |
| Node.js | 22.22.0 or newer |
| Lockfile | package-lock.json shipped inside the hajer package, with every transitive package and its integrity hash |
The version changes only with a hajer release, so one hajer version always runs one promptfoo version. To see
what your installation pins:
hajer eval --install-only
{
"detail": null,
"engineDir": "/home/you/.cache/hajer/engine/b0ce15c4d1973dfc",
"installed": true,
"lockfileDigest": "b0ce15c4d1973dfc997006c8871b25fe739859f8cc3e736c471a4536c46ae776",
"node": "/usr/local/bin/node",
"nodeFloor": "22.22.0",
"nodeMeetsFloor": true,
"nodeVersion": "22.23.2",
"npm": "/usr/local/bin/npm",
"pinnedVersion": "0.123.1",
"ready": true,
"reason": null
}
hajer doctor prints the same status without installing anything.
Installation
The first hajer eval (or hajer eval --install-only) installs the engine with
npm ci --omit=dev --no-audit --no-fund --loglevel=error into:
$HAJER_CACHE_DIR/engine/<first 16 hex characters of the lockfile's SHA-256>/
HAJER_CACHE_DIR defaults to $XDG_CACHE_HOME/hajer when XDG_CACHE_HOME is set, otherwise ~/.cache/hajer.
- Keyed by lockfile. A new
hajerversion with a new pin installs into a new directory beside the old one. Downgrading finds the old install still there. - Verified. After
npm ci, the engine must report the pinned version, and only then is a.readymarker written. A directory without the marker, or with a different version in it, is reinstalled. - Locked. The install holds a file lock (
<directory>.lock), so twohajer evalprocesses starting cold on the same machine do not install over each other; the second waits and reuses the result. The lock uses POSIXflock; on platforms without it the install proceeds unlocked. - Bounded. One install may take 600 s (
HAJER_EVAL_INSTALL_TIMEOUT_S). Raise it on a slow connection. - Online once. The install needs the npm registry. After that, runs need no registry access. In CI, cache
~/.cache/hajer/engine; see Running in CI.
Install failures exit before any suite runs:
| Exit code | Reason | Fix |
|---|---|---|
3 | NODE_MISSING, NODE_TOO_OLD, NPM_MISSING | Install Node.js 22.22.0 or newer, with npm, and put it first on PATH. |
4 | INSTALL_FAILED, INSTALL_TIMEOUT, SMOKE_FAILED | Check registry access, raise the timeout, or delete the engine directory named in the message and retry. |
The message names the reason and what to do:
hajer eval: NODE_TOO_OLD: Node.js 20.11.1 at /usr/bin/node is older than the 22.22.0 the eval engine needs; install Node.js 22.22.0 or newer (https://nodejs.org).
hajer eval --engine-help prints the pinned engine's promptfoo eval --help, installing it first if needed.
The engine's environment
promptfoo runs as a child process with your environment, with these changes:
| Variable | Value | Why |
|---|---|---|
HAJER_API_KEY | removed | The application under test cannot export to Hajer. Only the parent hajer eval process uploads. |
PROMPTFOO_DISABLE_TELEMETRY | 1 | promptfoo's product analytics. |
PROMPTFOO_DISABLE_UPDATE | 1 | promptfoo's version check. |
PROMPTFOO_DISABLE_SHARING | 1 | Result sharing to promptfoo's cloud. |
PROMPTFOO_DISABLE_REMOTE_GENERATION | 1 | Remote test generation. |
PROMPTFOO_DISABLE_REDTEAM_REMOTE_GENERATION | 1 | Remote red-team generation. |
PROMPTFOO_DISABLE_SHARE_EMAIL_REQUEST | 1 | The share-by-email prompt. |
NO_UPDATE_NOTIFIER | 1 | npm's update notifier. |
PROMPTFOO_CONFIG_DIR | <run dir>/promptfoo | promptfoo's results database and traces stay in the run directory, never in ~/.promptfoo. |
PROMPTFOO_CACHE_PATH | $HAJER_CACHE_DIR/promptfoo-cache | promptfoo's provider response cache, shared across runs. Pass --no-cache to bypass it. |
PROMPTFOO_PYTHON | the Python running hajer eval | The hook and your Python providers import the same hajer installation. |
HAJER_OTLP_ENDPOINT | http://127.0.0.1:<port> | Your application's spans go to promptfoo's local trace receiver. |
HAJER_ENVIRONMENT | eval | Every span says it came from an eval. |
HAJER_EVAL_* | the run's values | Read back by the hook and the provider binding. |
The trace receiver listens on 127.0.0.1, port 4318 by default (HAJER_EVAL_OTLP_PORT). If that port is busy,
a free port is used instead. hajer eval points only HAJER_OTLP_ENDPOINT at it and deliberately leaves the
generic OTEL_EXPORTER_OTLP_ENDPOINT alone: promptfoo's own OpenTelemetry SDK reads that variable as a complete
URL and would post to the receiver's root.
promptfoo exits 100 when any test fails or errors, and its own pass rate counts errored tests in the denominator.
hajer eval returns that code unchanged and reports failed and errored results separately in the payload.
hajer eval also adds a small configuration file of its own after yours, which turns on tracing and attaches the
beforeAll hook. It is written to the run directory; your files are not modified, and relative file:// paths
in your configuration resolve as they do with plain promptfoo.
That added configuration contributes an empty prompts list. promptfoo cannot merge a list with a map, so write
prompts in your configuration as a list (as every promptfoo example does), not as a label: prompt map.
Outbound connections
With the switches above, promptfoo itself still contacts three hosts on its own:
| Host | Why |
|---|---|
r.promptfoo.app | promptfoo sends one "telemetry disabled" event; this call is not behind its disable flag. |
169.254.169.254 | A cloud instance-metadata probe made at start-up by a bundled provider SDK. |
metadata.google.internal | The same, for Google Cloud. |
The two metadata probes come from SDKs bundled with promptfoo (the openai SDK's credential providers and
@azure/msal-common), and happen whatever providers your suite uses. None of these carries data about your suite,
none can be switched off without modifying promptfoo, and blocking them does not change the run. On a machine without
outbound access they are logged as unreachable and the run proceeds.
Everything else the run contacts is what your configuration asks for: your providers, your graders, and their
APIs. The upload to Hajer, when requested, is made by the parent hajer eval process after the engine exits.
Security
A promptfoo configuration is code. file:// providers, graders and hooks are Python or JavaScript that run on the
machine running hajer eval, with that machine's environment (minus HAJER_API_KEY). Nothing is sandboxed.
- Run evals from a pull request with the same trust you give its tests. On GitHub Actions, secrets are not exposed to pull requests from forks by default; keep it that way.
- Keep
HAJER_API_KEYand model provider keys in your CI's secret store, not inpromptfooconfig.yamlorhajer.yaml. - Outputs uploaded to Hajer are redacted with the client-side catalog first; see Redaction.