Skip to main content

Content capture and redaction

By default, model spans carry the messages sent to the model and the answers it returned, so you can read a conversation on the platform. Before any content is put on a span, the SDK redacts sensitive values in your process. This page covers what is captured, what is redacted, and how to change both.

Content capture​

With HAJER_CAPTURE_CONTENT=1 (the default), the SDK captures:

  • gen_ai.input.messages: the messages sent, including tool calls and tool results
  • gen_ai.output.messages: the answer, one assistant message per choice, with its finish reason
  • gen_ai.system_instructions: an out-of-band system prompt (Anthropic system, Responses instructions, google-genai config.system_instruction)
  • gen_ai.tool.call.arguments on hajer.tool spans

Messages use the OpenTelemetry GenAI message format, whichever library made the call:

[
{"role": "user", "parts": [{"type": "text", "content": "where is my order"}]},
{"role": "assistant", "parts": [{"type": "tool_call", "id": "call-1", "name": "lookup", "arguments": {"id": "o-1"}}]},
{"role": "tool", "parts": [{"type": "tool_call_response", "id": "call-1", "result": "shipped"}]}
]

How each library's shapes map onto this format:

  • OpenAI chat messages, Responses items (function_call, function_call_output), Anthropic blocks (tool_use, tool_result) and google-genai parts (function_call, function_response) all become text, tool_call and tool_call_response parts. Tool-call arguments a provider sends as a JSON string are parsed into a document when they are valid JSON.
  • Roles are folded onto user, assistant, system and tool: human becomes user, ai and model become assistant, developer becomes system, and function becomes tool.
  • Non-text parts, such as images, documents and audio, are recorded as a part naming their type, without the data.
  • A message in a shape the SDK does not recognise is kept as the text of its JSON rather than dropped.
  • The output is one assistant message per choice, with the finish reason on it and any tool calls on the first. A streamed call's output is the text as it arrived.
  • gen_ai.system_instructions holds only an out-of-band system prompt. OpenAI chat keeps its system message in gen_ai.input.messages, where you passed it.

Embedding calls never record the input or the vectors.

To turn content capture off:

export HAJER_CAPTURE_CONTENT=0

With capture off, spans keep the model, request settings, declared tool names, finish reasons, timing and token usage, and carry no text. Tool spans carry no arguments.

Size limits​

HAJER_WRAPPED_CALL_MAX_BYTES (default 32768) bounds the content of one call. The input, output and system instructions of a span together must fit, filled in that order. When they do not fit, the longest texts are clipped in the middle: the head and tail are kept, with a marker between them giving the full length and the SHA-256 of the whole text. The span's hajer.limitations says that content was clipped. If even the structure of a document does not fit, the attribute is left out and hajer.limitations says so.

Tool arguments on hajer.tool spans are clipped to the same bound.

HAJER_BODY_MAX_BYTES (default 65536) is separate: it bounds how much of a raw response body (with_raw_response) is buffered to read the answer out of it.

Client-side redaction​

With HAJER_REDACT_CLIENT=1 (the default), the SDK replaces sensitive values with [redacted:<CATEGORY>] in every message, answer, tool argument and tool result, before a span carries it. The value never leaves your process.

The rules come from the Hajer platform's redaction catalog, bundled with the SDK. hajer doctor prints the catalog version (redaction-rules@6 in SDK 0.2.0). These categories are redacted by default:

CategoryWhat it matches
ANTHROPIC_KEY, OPENAI_KEY, HAJER_KEY, AWS_ACCESS_KEY_ID, BEARER_TOKENAPI keys and bearer tokens
CARDPayment card numbers (Luhn-checked)
IBANInternational bank account numbers (mod-97-checked)
VINVehicle identification numbers (check-digit-checked)
US_SSNUS Social Security numbers
NATIONAL_IDCanadian SIN, Brazilian CPF, Indian Aadhaar, French NIR and UK National Insurance numbers
EMAILEmail addresses
PHONEPhone numbers
PERSON_NAMEPerson names that follow an honorific (Mr., Dr., Professor and similar)
POSTAL_ADDRESSStreet addresses that end in a postcode
DATE_OF_BIRTHA date introduced by a phrase such as "date of birth", "born on" or "DOB"
ACCOUNT_LIKELong digit runs that look like account numbers (12 digits or more)

The rules marked as checked confirm a match with its checksum before redacting it, to avoid false positives on arbitrary numbers. The name, address and date-of-birth rules match only the shapes described, so a bare name or date elsewhere in a text is not redacted; add an extra_rules pattern if you need more.

In addition:

  • Secret field names. A value whose key looks like a secret is replaced with [redacted:FIELD], whatever it holds. This covers keys named pin, cvv, otp, pwd and similar, and keys ending in password, secret, token, apikey, authorization, cookie or credential(s), among others. Keys are compared ignoring case and punctuation, so client_secret, x-api-key and session_token are withheld, while max_tokens is not.
  • Budgets. The redaction pass is bounded: 1 MiB per document, 50,000 values, 64 levels of nesting and 65,536 characters per value. Past a budget, it stops scanning and replaces what it did not read with [redacted:UNSCANNED]. It never sends a value it did not scan.
  • The pass never raises into your code.

For example, a tool span declared with arguments={"email": "[email protected]", "id": 5, "api_key": "zzz"} is exported with:

{"email": "[redacted:EMAIL]", "id": 5, "api_key": "[redacted:FIELD]"}

Custom policies​

Use hajer.build_policy to adjust the rules, and hajer.configure(policy=...) to apply them to every span from then on:

import hajer

policy = hajer.build_policy(
classes_off=["ACCOUNT_LIKE"], # a category your team has decided to keep
extra_rules=[("CUSTOMER_REF", r"\bCUS-\d{6}\b")], # your own patterns: (category, regex)
paths_exempt=["request.account"], # paths that must travel unredacted
)
hajer.configure(policy=policy)
hajer.build_policy(*, classes_off=(), extra_rules=(), paths_exempt=()) -> ClientRedactionPolicy
  • classes_off: categories from the table above to turn off.
  • extra_rules: (category, pattern) pairs. A match is replaced with [redacted:<category>].
  • paths_exempt: paths in the document to leave untouched.

build_policy validates everything when you call it: a pattern that does not compile raises re.error, and a path it cannot parse raises a ValueError. Both happen where you build the policy, never during a model call.

note

configure(policy=...) replaces the whole configuration, including any tracer_provider or settings you passed before. Pass all three together if you use more than one.

Redacting a document yourself​

hajer.redact_document runs the same pass on any JSON-compatible value, for example to check a policy in a test:

redacted, entries = hajer.redact_document(
{"note": "Card 4111 1111 1111 1111, mail [email protected]"},
policy=hajer.build_policy(),
)
print(redacted)
for entry in entries:
print(entry.category, entry.path, entry.count)
{'note': 'Card [redacted:CARD], mail [redacted:EMAIL]'}
EMAIL note 1
CARD note 1

It returns the redacted copy and a tuple of RedactionEntry objects (category, path, count), one per category found at each path. It never raises.

Turning redaction off​

export HAJER_REDACT_CLIENT=0

This sends content as captured. Prefer paths_exempt or classes_off if you need only specific values to travel.

note

Redaction applies to exported spans. Calls you read in-process through hajer.wrapped_calls() or hajer.scope() hold the content as captured.

Call-site capture​

With HAJER_CAPTURE_CALL_SITE=1 (the default), each model span records where your code made the call:

  • code.function.name, code.file.path, code.line.number: the innermost application frame
  • server.address, server.port: the provider endpoint

The local call record (WrappedCall.caller_frames, see scope) keeps up to 8 application frames, innermost first, each with its module, qualified name, file and line. An application frame is any frame whose file is not inside the SDK, the Python standard library or an installed-package directory. An application installed into site-packages without an editable install therefore has no application frames: the call is recorded without them, and hajer.limitations says CALLER_FRAMES_UNRESOLVED.

File paths are always relative to the project root: HAJER_PROJECT_ROOT, or the working directory if it is not set. A file outside the root is recorded as <outside>/<file name>, and a home directory is never part of a recorded path. These are locations, not values, so they are captured independently of content capture.

If your process starts from a directory other than your checkout, such as / under systemd, set the root so paths are useful:

export HAJER_PROJECT_ROOT=/srv/support-api

To record neither the call site nor the endpoint:

export HAJER_CAPTURE_CALL_SITE=0

Outbound HTTP requests​

With HAJER_CAPTURE_HTTP=1 (default 0), the SDK also records your application's other outbound requests made through httpx (and httpx2) default transports, such as a search API or a data vendor, as HTTP client spans. The transports are patched at the first wrap, instrument or attach call.

Each span is named after the HTTP method and carries:

  • http.request.method
  • url.template: the scheme, host and path, with sensitive path segments replaced. A numeric segment becomes {id}, a segment the redaction catalog recognises becomes {redacted}, and a long random-looking segment becomes {token}.
  • http.response.status_code
  • server.address, server.port and the call site

Headers, query strings and bodies are never recorded. Requests made by a wrapped provider client and the SDK's own requests are not recorded as HTTP spans.

hajer.uninstrument() removes the transport patches.