Skip to main content

Evals quickstart

hajer eval runs ordinary promptfoo configurations on a pinned copy of promptfoo, and reports each run to Hajer against the workflows and obligations its tests cover. Nothing about promptfoo's configuration format changes. You add:

  • a hajer.yaml at the repository root naming your suites and obligations,
  • a metadata.hajer block on the tests that cover a workflow,
  • a one-line decorator on your promptfoo Python provider.

This page takes you from nothing to a run on the Tests page.

Prerequisites​

  • Python 3.11 or newer.
  • Node.js 22.22.0 or newer, with npm, on your PATH. The engine installs itself with npm ci on the first run; it needs the npm registry once.
  • For uploads: a Hajer team API key and team id, and the repository connected to Hajer through GitHub.
node --version # v22.22.0 or newer
npm --version

1. Install​

pip install "hajer[evals]"

The evals extra includes the OpenTelemetry exporter (hajer[otel]) and a YAML reader. The hajer command is also available as python -m hajer.

2. Instrument the code under test​

Mark the workflow your tests exercise. These are the same decorators your production code uses; see Workflows, components and tools.

app.py
import hajer

@hajer.tool("tool_refund_status", name="get_refund_status")
def get_refund_status(customer_id: str) -> str:
return {"cust_123": "pending"}.get(customer_id, "unknown")

@hajer.workflow("wf_support")
def handle(message: str, customer_id: str) -> str:
status = get_refund_status(customer_id)
if status == "pending":
return "Your refund is still pending."
return "I cannot find a refund for this account."

3. Write the provider​

promptfoo reaches your application through a provider. @hajer.evals.provider ties the spans your code emits to the test that triggered them.

provider.py
import hajer.evals
from app import handle

@hajer.evals.provider
def call_api(prompt, options, context):
variables = context.get("vars", {})
return {"output": handle(variables["message"], variables["customer_id"])}

4. Write the promptfoo configuration​

promptfooconfig.yaml
description: Customer support regression

providers:
- id: file://provider.py
label: Support agent

prompts:
- "{{message}}"

defaultTest:
metadata:
hajer:
workflowId: wf_support # every test inherits this

tests:
- description: Pending refund is reported as pending
metadata:
testCaseId: eval_refund_pending
hajer:
obligationIds: [obl_refund_status_disclosed]
vars:
message: Has my refund gone through?
customer_id: cust_123
assert:
- type: contains
value: pending
- type: trajectory:tool-used
value: get_refund_status

5. Declare the suite in hajer.yaml​

Put this file at the repository root:

hajer.yaml
version: 1
suites:
- id: support
path: promptfooconfig.yaml
obligations:
- id: obl_refund_status_disclosed
title: Refund status is disclosed accurately
workflow: wf_support

The full schema is in the hajer.yaml reference.

6. Run it​

hajer eval

The first run installs the pinned engine into ~/.cache/hajer/engine/. Each suite ends with one summary line:

hajer eval (suite support): passed: 1/1 passed, 0 failed, 0 errored; run evalrun_c4bc72e7...; payload /home/you/.cache/hajer/runs/evalrun_c4bc72e7.../payload.json; upload not requested

The exit code is 0 when every test passed and 100 when any failed. See the CLI reference for the rest.

7. Upload the run​

  1. In the Hajer app, open Settings → API keys and create a key. The dialog shows the key once, together with the team id and base URL.
  2. In Settings → GitHub, connect the Hajer GitHub App and use Connect repository to link the repository that holds hajer.yaml. Linking scans the repository, so its suites and obligations appear on the Tests page before the first upload.
  3. Run with credentials:
export HAJER_API_KEY=...
export HAJER_TEAM_ID=...
hajer eval --upload

The summary line now ends in upload uploaded. The run appears on the Tests page of the repository named by the commit's git context.

note

An upload that fails, or is skipped because credentials are missing, never changes the exit code. A green eval with a failed upload is a green eval. The payload stays on disk; see Uploading runs.

Next steps​