Uploading runs
hajer eval --upload sends each suite's payload to Hajer once the engine has finished. The platform attaches the
run to the linked GitHub repository named by the run's git context, and it appears on that repository's Tests
page.
Before you upload
- Create a team API key in Settings → API keys. Set
HAJER_API_KEYandHAJER_TEAM_IDfrom the values shown. - Connect the Hajer GitHub App in Settings → GitHub and link the repository with Connect repository.
export HAJER_API_KEY=...
export HAJER_TEAM_ID=...
# export HAJER_BASE_URL=https://api.hajer.ai # the default; change only for a self-hosted platform
hajer eval --upload
Without HAJER_API_KEY and HAJER_TEAM_ID, or with HAJER_DISABLED=1, nothing is sent and the upload is reported
skipped (INERT). A pull request from a fork, whose CI has no secrets, still runs its evals and passes or fails on
their results.
The request
POST {HAJER_BASE_URL}/api/teams/{HAJER_TEAM_ID}/eval-runs
Authorization: Bearer <HAJER_API_KEY>
Content-Type: application/json
Idempotency-Key: <runId>
The body is the payload as compact JSON with sorted keys. The run id is the idempotency key on every attempt, so retrying the same upload stores the run once.
Retries
| Response | What hajer eval does |
|---|---|
2xx | uploaded. |
409 Conflict | Treated as uploaded: a run with this id is already stored. |
any other 4xx | failed (REFUSED). Not retried; the same body would get the same answer. |
5xx, connection error | Retried. After the last attempt, failed (UNREACHABLE). |
| timeout | Retried. After the last attempt, failed (TIMEOUT). |
There are 3 attempts by default (HAJER_EVAL_UPLOAD_ATTEMPTS). The wait between attempts starts at 200 ms and
doubles, capped at 30 s (HAJER_EVAL_UPLOAD_BACKOFF_INITIAL_MS, HAJER_EVAL_UPLOAD_BACKOFF_MAX_MS). Each request
has a 30 s deadline (HAJER_EVAL_UPLOAD_DEADLINE_MS).
On the platform side, re-sending the identical payload answers 200, and a first upload answers 201. Sending a
different payload under a run id that is already stored answers 409, which hajer eval reports as uploaded;
the platform keeps the first one.
Outcomes
The summary line ends with upload <outcome>:
| Outcome | Meaning |
|---|---|
not requested | --upload was not given, or --no-upload was. |
uploaded | Stored (or already stored). |
skipped (INERT) | No credentials, or HAJER_DISABLED=1. Nothing was sent. |
failed (BODY_OVER_BOUND) | Still over the size limit after trimming. Nothing was sent. |
failed (REFUSED) | The platform answered with a 4xx other than 409. Check the key, the team id and the base URL. |
failed (UNREACHABLE) | No usable answer after every attempt. |
failed (TIMEOUT) | The last attempt ran out of time. |
failed (NO_RUN_ID) | Internal error: the payload had no run id. |
Whatever the outcome, the exit code is the engine's, and the payload stays at
$HAJER_CACHE_DIR/runs/<run id>/payload.json.
Size limit and trimming
The body is limited to 8,388,608 bytes (8 MiB, HAJER_EVAL_UPLOAD_MAX_BYTES). If the encoded payload is larger,
hajer eval trims it in this order, re-checking after each step:
- Remove
spansfrom every result (thespanSummaryis kept). - Set
outputtonullon every result. - If it is still too large, send nothing:
failed (BODY_OVER_BOUND).
The platform enforces the same 8 MiB limit, so raising HAJER_EVAL_UPLOAD_MAX_BYTES past it turns a local refusal
into a 413 from the platform (failed (REFUSED)). To make large runs fit, lower HAJER_EVAL_SPANS_MAX or
HAJER_EVAL_OUTPUT_MAX_CHARS instead.
Where the run lands
The payload names no project or repository explicitly. The platform matches it against the team's linked repositories using, in order:
git.ci.repository, when it is anowner/name(GitHub Actions sets this fromGITHUB_REPOSITORY),git.remoteUrl, when it is agithub.comremote (HTTPS, SSH or[email protected]:owner/name.git).
Matching is case-insensitive. A run that matches no linked repository is still stored, as an unlinked run of the team. The platform also records whether the run's branch is the repository's default branch.
Git and CI context
Before the engine starts, hajer eval reads the commit context from git in the suite's directory and from a
fixed list of CI variables. Nothing here can fail the run: any value it cannot read is null. Each git call is
limited to 5 s (HAJER_EVAL_GIT_TIMEOUT_S).
From git:
commitSha:git rev-parse HEADbranch:git rev-parse --abbrev-ref HEAD(nullon a detached head)dirty: whethergit status --porcelain --untracked-files=noreports changesremoteUrl:remote.origin.url, with credentials removed
From CI, the provider is detected by its own flag, in this order:
Provider (git.ci.provider) | Detected by | Variables read |
|---|---|---|
github-actions | GITHUB_ACTIONS | GITHUB_SHA, GITHUB_REF, GITHUB_REF_NAME, GITHUB_HEAD_REF, GITHUB_BASE_REF, GITHUB_RUN_ID, GITHUB_REPOSITORY, GITHUB_SERVER_URL |
gitlab-ci | GITLAB_CI | CI_COMMIT_SHA, CI_COMMIT_REF_NAME, CI_MERGE_REQUEST_IID, CI_MERGE_REQUEST_SOURCE_BRANCH_NAME, CI_MERGE_REQUEST_TARGET_BRANCH_NAME, CI_PIPELINE_ID, CI_PROJECT_URL |
circleci | CIRCLECI | CIRCLE_SHA1, CIRCLE_BRANCH, CIRCLE_PULL_REQUEST, CIRCLE_BUILD_NUM, CIRCLE_REPOSITORY_URL |
buildkite | BUILDKITE | BUILDKITE_COMMIT, BUILDKITE_BRANCH, BUILDKITE_PULL_REQUEST, BUILDKITE_PULL_REQUEST_BASE_BRANCH, BUILDKITE_BUILD_ID, BUILDKITE_REPO |
generic | CI | none |
These variables (plus CI) are the only environment variables read for the context; your CI secrets are never
read. When a CI provider supplies a commit and branch, they take precedence over the local git values, because
CI often checks out a detached merge commit. For a GitHub pull request, the PR number comes from GITHUB_REF
(refs/pull/<n>/merge) and the branch from GITHUB_HEAD_REF.
Credential stripping
Every repository URL, local or from CI, has its credentials removed before it is stored:
| Input | Stored as |
|---|---|
https://x-access-token:[email protected]/acme/app.git | https://github.com/acme/app.git |
https://[email protected]/acme/app.git | https://github.com/acme/app.git |
ssh://git:[email protected]/acme/app.git | ssh://[email protected]/acme/app.git |
[email protected]:acme/app.git | unchanged |
The payload
One JSON document per suite run, schemaVersion: 1. It is written to payload.json whether or not you upload,
and is the exact upload body. Keys are camelCase.
Run fields
| Field | Description |
|---|---|
schemaVersion | Always 1. |
runId | evalrun_ plus 32 hex characters, created before the engine starts. |
createdAt | ISO 8601 UTC timestamp of the run's start. |
status | passed, failed, errored, or aborted (no results and an engine exit code other than 0 or 100). Otherwise the worst result outcome. |
engineExitCode | promptfoo's exit code. |
engine | name (promptfoo), version, lockfileDigest, nodeVersion, evalId. |
sdk | version of the hajer package. |
git | commitSha, branch, dirty, remoteUrl, and ci with provider, runId, prNumber, baseRef, headRef, repository. null outside a git checkout and outside CI. |
suiteId | The suite's id in hajer.yaml; null for a -c run. |
config | path (relative to hajer.yaml for a declared suite), description, providerIds. |
filters | workflowId and obligationIds from --workflow and --obligation. |
stats | total, passed, failed, errored, durationMs, tokenUsage, cost. |
warnings | Metadata warnings such as W_NO_TEST_CASE_ID. |
results | One entry per test and prompt. |
Result fields
| Field | Description |
|---|---|
testCaseId | The test's id: from its trace, else metadata.testCaseId, else <testIdx>-<promptIdx>. |
testIdx, promptIdx | promptfoo's indices. |
correlation | platform when the test has a workflowId, else none. |
workflowId, obligationIds, componentIds, sourceTraceIds, generatedBy, provenance | From the test's effective metadata.hajer. |
description | The test's description. |
provider | id and label. |
outcome | passed; errored when the provider produced no output to grade; otherwise failed. |
score, error | promptfoo's score and error message. |
assertions | Each with type, metric, passed, score, reason, weight. |
latencyMs, tokenUsage, cost | As promptfoo reported them. |
traceId, evaluationId | promptfoo's trace and evaluation ids. |
output | The provider's output, redacted, then clipped to HAJER_EVAL_OUTPUT_MAX_CHARS. |
spans | Up to HAJER_EVAL_SPANS_MAX spans of the test's trace, earliest first: spanId, parentSpanId, name, startTime, endTime (epoch milliseconds), status, attributes. |
spanSummary | Over all spans, before the cap: count, errorCount, toolNames, componentIds, workflowIds. |
Span attributes are limited to these prefixes: hajer., gen_ai., deployment., tool., session., user.,
server., error., code.. Everything else a span carries (http.*, db.*, framework attributes) stays on
your machine.
Outputs are redacted with the same client-side catalog the tracing SDK uses before they are clipped, so a clip
never exposes part of a value the redaction would have caught. HAJER_REDACT_CLIENT=0 turns this off. See
Redaction.
Settings
| Variable | Default | Purpose |
|---|---|---|
HAJER_API_KEY | none | Team API key. Required to upload. |
HAJER_TEAM_ID | none | Team id. Required to upload. |
HAJER_BASE_URL | https://api.hajer.ai | The platform. Change only for a self-hosted platform. |
HAJER_DISABLED | 0 | 1 skips the upload. |
HAJER_CACHE_DIR | $XDG_CACHE_HOME/hajer, else ~/.cache/hajer | Engine installs and run directories. |
HAJER_EVAL_RUNS_KEEP | 20 | Run directories kept in the cache. |
HAJER_EVAL_UPLOAD_ATTEMPTS | 3 | Attempts per upload. |
HAJER_EVAL_UPLOAD_BACKOFF_INITIAL_MS | 200 | First wait between attempts. |
HAJER_EVAL_UPLOAD_BACKOFF_MAX_MS | 30000 | Ceiling of the doubling wait. |
HAJER_EVAL_UPLOAD_DEADLINE_MS | 30000 | Deadline of one upload request. |
HAJER_EVAL_UPLOAD_MAX_BYTES | 8388608 | Payload size limit before trimming. The platform also enforces 8 MiB. |
HAJER_EVAL_OUTPUT_MAX_CHARS | 4096 | Characters of each output kept. |
HAJER_EVAL_SPANS_MAX | 256 | Spans per result kept; the summary counts all. |
HAJER_EVAL_GIT_TIMEOUT_S | 5 | Limit on each git read. |
HAJER_EVAL_INSTALL_TIMEOUT_S | 600 | Limit on the engine install. See The engine. |
HAJER_EVAL_OTLP_PORT | 4318 | Local trace receiver port; a busy port is replaced by a free one. |
HAJER_TRACE_FLUSH_TIMEOUT_MS | 2000 | How long a provider waits for its spans to reach the receiver. |
Numeric settings must be positive integers. HAJER_EVAL_RUN_ID, HAJER_EVAL_WORKFLOW, HAJER_EVAL_OBLIGATIONS,
HAJER_EVAL_HOOK_REPORT, HAJER_EVAL_MANIFEST and HAJER_EVAL_MANIFEST_OBLIGATIONS are set by hajer eval for
the engine process; do not set them yourself. For the rest of the SDK's settings, see
Configuration.