Skip to main content

Providers

promptfoo calls your application through a provider. For a Python provider, hajer.evals adds one decorator, @hajer.evals.provider, and one context manager, hajer.evals.bind(context). Either one connects the spans your application emits to the promptfoo test that triggered them. promptfoo's trace-based assertions (trajectory:*, trace-span-count) can then read those spans, and Hajer can join each result to its trace.

The decorator​

provider.py
import hajer.evals
from app import handle

@hajer.evals.provider
def call_api(prompt, options, context):
variables = context.get("vars", {})
message = variables.get("message", prompt)
customer_id = variables.get("customer_id", "")
return {"output": handle(str(message), str(customer_id))}

Reference it from the configuration as usual:

promptfooconfig.yaml
providers:
- id: file://provider.py
label: Support agent

promptfoo runs the provider with its own directory on sys.path, so from app import handle resolves against files next to provider.py.

The decorator keeps the function's name, signature and return value. It reads the engine's context from the third positional argument or from a context= keyword. Called without a context, the function still runs, unbound.

Async providers​

async def works the same way:

provider.py
import hajer.evals
from app import handle_async

@hajer.evals.provider
async def call_api(prompt, options, context):
variables = context.get("vars", {})
return {"output": await handle_async(variables["message"], variables["customer_id"])}

The context manager​

Use hajer.evals.bind(context) when only part of the provider should belong to the test, or when the provider is not a single function:

provider.py
import hajer.evals
from app import build_agent

def call_api(prompt, options, context):
agent = build_agent(options.get("config", {})) # setup, outside the test's trace
with hajer.evals.bind(context):
output = agent.run(prompt)
return {"output": output}

What the binding does​

On entry, it reads from the engine's context:

context fieldUsed for
traceparentThe W3C trace context promptfoo opened for this test.
testCaseIdThe test's id.
test.metadata.hajer.workflowIdThe workflow the test targets.
test.metadata.hajer.obligationIdsThe obligations the test covers.

Inside the block:

  • A span opened while no other span is active becomes a child of traceparent. Your workflow span, and everything under it, lands in the trace promptfoo grades.
  • Every span carries hajer.eval.run.id, hajer.eval.test_case.id, hajer.eval.workflow.id and hajer.eval.obligation.ids where known. The run id comes from HAJER_EVAL_RUN_ID, which hajer eval sets.

On exit, it flushes the spans, waiting up to HAJER_TRACE_FLUSH_TIMEOUT_MS (default 2000), so they reach the engine's trace receiver before the test's assertions run.

Missing or mistyped context fields are treated as absent. The binding never raises inside your provider.

Emit spans the assertions can read​

The binding only connects spans; your code still has to emit them. Use the same hajer.workflow, hajer.component and hajer.tool declarations as in production (Workflows, components and tools):

app.py
import hajer

@hajer.tool("tool_refund_status", name="get_refund_status")
def get_refund_status(customer_id: str) -> str:
return {"cust_123": "pending", "cust_456": "completed"}.get(customer_id, "unknown")

@hajer.component("cmp_refund_agent")
def refund_agent(customer_id: str) -> str:
status = get_refund_status(customer_id)
if status == "pending":
return "Your refund is still pending: it has been approved but has not completed yet."
if status == "completed":
return "Your refund has completed."
return "I cannot find a refund for this account."

@hajer.workflow("wf_support")
def handle(message: str, customer_id: str) -> str:
if "refund" in message.lower():
return refund_agent(customer_id)
return "I can help with refund questions."

What each span gives the assertions:

DeclarationSpan nameAttributes assertions use
hajer.workflow("wf_support")workflow wf_supporthajer.workflow.id
hajer.component("cmp_refund_agent")component cmp_refund_agenthajer.component.id
hajer.tool("tool_refund_status", name="get_refund_status")execute_tool get_refund_statusgen_ai.tool.name, hajer.tool.id

promptfoo's trajectory:tool-used matches on the tool name, which is name= (or the tool id when you omit it):

assert:
- type: trajectory:tool-used
value: get_refund_status
- type: trace-span-count
value:
pattern: "workflow *"
min: 1

Model calls recorded through hajer.wrap(client) or hajer.instrument() inside the workflow appear as gen_ai spans in the same trace. See Instrumenting clients.

How the application runs under hajer eval​

  • Spans export to the engine's local trace receiver on 127.0.0.1, not to Hajer. hajer eval points HAJER_OTLP_ENDPOINT there and removes HAJER_API_KEY from the engine's environment, so test traffic is never exported as production traces. The run's results, including a capped copy of each test's spans, reach Hajer only through hajer eval --upload.
  • Every span carries deployment.environment.name (and deployment.environment) set to eval.
  • The emitter needs the OpenTelemetry SDK, which hajer[evals] installs. Without it, the declarations still run the code they wrap but emit nothing, and warn once.
  • If the process already has an OpenTelemetry tracer provider configured, the SDK uses it and its exporters.