Providers
promptfoo calls your application through a provider. For a
Python provider, hajer.evals adds one decorator, @hajer.evals.provider, and one context manager,
hajer.evals.bind(context). Either one connects the spans your application emits to the promptfoo test that
triggered them. promptfoo's trace-based assertions (trajectory:*, trace-span-count) can then read those spans,
and Hajer can join each result to its trace.
The decorator
import hajer.evals
from app import handle
@hajer.evals.provider
def call_api(prompt, options, context):
variables = context.get("vars", {})
message = variables.get("message", prompt)
customer_id = variables.get("customer_id", "")
return {"output": handle(str(message), str(customer_id))}
Reference it from the configuration as usual:
providers:
- id: file://provider.py
label: Support agent
promptfoo runs the provider with its own directory on sys.path, so from app import handle resolves against
files next to provider.py.
The decorator keeps the function's name, signature and return value. It reads the engine's context from the
third positional argument or from a context= keyword. Called without a context, the function still runs,
unbound.
Async providers
async def works the same way:
import hajer.evals
from app import handle_async
@hajer.evals.provider
async def call_api(prompt, options, context):
variables = context.get("vars", {})
return {"output": await handle_async(variables["message"], variables["customer_id"])}
The context manager
Use hajer.evals.bind(context) when only part of the provider should belong to the test, or when the provider is
not a single function:
import hajer.evals
from app import build_agent
def call_api(prompt, options, context):
agent = build_agent(options.get("config", {})) # setup, outside the test's trace
with hajer.evals.bind(context):
output = agent.run(prompt)
return {"output": output}
What the binding does
On entry, it reads from the engine's context:
context field | Used for |
|---|---|
traceparent | The W3C trace context promptfoo opened for this test. |
testCaseId | The test's id. |
test.metadata.hajer.workflowId | The workflow the test targets. |
test.metadata.hajer.obligationIds | The obligations the test covers. |
Inside the block:
- A span opened while no other span is active becomes a child of
traceparent. Your workflow span, and everything under it, lands in the trace promptfoo grades. - Every span carries
hajer.eval.run.id,hajer.eval.test_case.id,hajer.eval.workflow.idandhajer.eval.obligation.idswhere known. The run id comes fromHAJER_EVAL_RUN_ID, whichhajer evalsets.
On exit, it flushes the spans, waiting up to HAJER_TRACE_FLUSH_TIMEOUT_MS (default 2000), so they reach the
engine's trace receiver before the test's assertions run.
Missing or mistyped context fields are treated as absent. The binding never raises inside your provider.
Emit spans the assertions can read
The binding only connects spans; your code still has to emit them. Use the same hajer.workflow,
hajer.component and hajer.tool declarations as in production
(Workflows, components and tools):
import hajer
@hajer.tool("tool_refund_status", name="get_refund_status")
def get_refund_status(customer_id: str) -> str:
return {"cust_123": "pending", "cust_456": "completed"}.get(customer_id, "unknown")
@hajer.component("cmp_refund_agent")
def refund_agent(customer_id: str) -> str:
status = get_refund_status(customer_id)
if status == "pending":
return "Your refund is still pending: it has been approved but has not completed yet."
if status == "completed":
return "Your refund has completed."
return "I cannot find a refund for this account."
@hajer.workflow("wf_support")
def handle(message: str, customer_id: str) -> str:
if "refund" in message.lower():
return refund_agent(customer_id)
return "I can help with refund questions."
What each span gives the assertions:
| Declaration | Span name | Attributes assertions use |
|---|---|---|
hajer.workflow("wf_support") | workflow wf_support | hajer.workflow.id |
hajer.component("cmp_refund_agent") | component cmp_refund_agent | hajer.component.id |
hajer.tool("tool_refund_status", name="get_refund_status") | execute_tool get_refund_status | gen_ai.tool.name, hajer.tool.id |
promptfoo's trajectory:tool-used matches on the tool name, which is name= (or the tool id when you omit it):
assert:
- type: trajectory:tool-used
value: get_refund_status
- type: trace-span-count
value:
pattern: "workflow *"
min: 1
Model calls recorded through hajer.wrap(client) or hajer.instrument() inside the workflow appear as gen_ai
spans in the same trace. See Instrumenting clients.
How the application runs under hajer eval
- Spans export to the engine's local trace receiver on
127.0.0.1, not to Hajer.hajer evalpointsHAJER_OTLP_ENDPOINTthere and removesHAJER_API_KEYfrom the engine's environment, so test traffic is never exported as production traces. The run's results, including a capped copy of each test's spans, reach Hajer only throughhajer eval --upload. - Every span carries
deployment.environment.name(anddeployment.environment) set toeval. - The emitter needs the OpenTelemetry SDK, which
hajer[evals]installs. Without it, the declarations still run the code they wrap but emit nothing, and warn once. - If the process already has an OpenTelemetry tracer provider configured, the SDK uses it and its exporters.