Instrumenting clients
The SDK records a model call at the provider client that makes it. Each recorded call becomes one gen_ai span (a generation on the platform) when the call settles. There are three ways to instrument clients.
| Approach | What it instruments | Code change | Use it when |
|---|---|---|---|
hajer.wrap(client) | One client object you pass in | One line where you build the client | You build your clients in a few known places. |
hajer.instrument() | Every openai and anthropic client built afterwards | One line at startup | You build clients in many places, or inside libraries. |
Attach mode (hajer.attach()) | Every client of every supported library built afterwards | None, or one import | You want telemetry from an application without editing it. |
All three record calls the same way, and you only need one of them.
wrap: one client
import hajer
from openai import AsyncOpenAI, OpenAI
from anthropic import Anthropic
openai_client = hajer.wrap(OpenAI())
async_client = hajer.wrap(AsyncOpenAI())
anthropic_client = hajer.wrap(Anthropic())
wrap instruments the object in place and returns it. Calling it twice on the same client has no effect.
wrap also accepts a google.genai.Client and the litellm module:
import litellm
from google import genai
import hajer
gemini = hajer.wrap(genai.Client())
hajer.wrap(litellm) # patches litellm.completion and litellm.acompletion
If wrap does not recognise the object, it raises hajer.UnsupportedClientError rather than silently recording nothing. A unittest.mock object is the exception: it is returned unchanged and records nothing, so tests that patch the provider class keep working.
LangChain chat models
Wrap the chat model itself. The SDK instruments the provider client the chat model holds and returns the same chat model, so you can pass it anywhere that accepts a BaseChatModel, including after bind_tools.
from langchain_openai import ChatOpenAI
from langchain_anthropic import ChatAnthropic
import hajer
llm = hajer.wrap(ChatOpenAI(model="gpt-5"))
claude = hajer.wrap(ChatAnthropic(model="claude-sonnet-5"))
reply = llm.invoke("Summarize this ticket")
The span records the request that left your process, not LangChain's internal representation of it. For ChatOpenAI the SDK instruments root_client and root_async_client, including the with_raw_response calls LangChain makes; for ChatAnthropic, _client and _async_client. A chat model that keeps its provider client anywhere else raises UnsupportedClientError rather than recording nothing.
instrument(): every client built from now on
import hajer
hajer.instrument() # call once, before you build clients
from openai import OpenAI
client = OpenAI() # instrumented at construction
instrument() patches the constructors of these classes so every instance is instrumented when it is built:
openai:OpenAI,AsyncOpenAI,AzureOpenAI,AsyncAzureOpenAIanthropic:Anthropic,AsyncAnthropic,AnthropicBedrock,AsyncAnthropicBedrock,AnthropicVertex,AsyncAnthropicVertex
This also covers LangChain chat models, because they build these clients internally. If the library is not imported yet, instrument() installs an import hook and patches it when it is imported.
instrument() returns an Instrumentation receipt. A second call returns the same receipt. hajer.uninstrument() restores the original constructors and removes the import hook.
hajer.instrument()
...
hajer.uninstrument()
instrument() does not cover google-genai or litellm. Use hajer.wrap(...) or attach mode for those.
Attach mode: no code change
Attach mode instruments every library in the support table for the whole process. It is opt-in. There are three ways to turn it on, and they all run the same hook.
From the shell, with no change to your code:
PYTHONPATH="$(python -m hajer attach-path)" HAJER_ATTACH=1 python -m your_app
hajer attach-path prints a directory containing a sitecustomize.py that Python imports at startup. Python imports only the first sitecustomize it finds, so if your environment already has one, this would shadow it. In that case use the one-line import below instead.
With one import in your entry point:
import hajer.autoattach # noqa: F401 attaches only when HAJER_ATTACH is on
This line reads HAJER_ATTACH, so you can leave it in your code and use the variable as the switch.
In code, unconditionally:
import hajer
attachment = hajer.attach()
print(attachment.describe())
# hajer attached to 10 client classes (openai, anthropic); exporting every model call as a span
hajer.attach() does not read HAJER_ATTACH. It is idempotent: a second call returns the standing attachment. It returns an Attachment with:
classes: the client classes it patchedmodules: the libraries it foundinert:Truewhen no key is configured, so nothing will be sentdescribe(): a one-line summary for a log
Libraries that are already imported are patched immediately. Libraries imported later are patched by an import hook.
hajer.attachment() returns the current Attachment, or None when the process is not attached.
hajer.detach() removes the import hook and stops instrumenting clients built afterwards. It does not un-patch anything: clients that are already instrumented keep recording and exporting. To stop a process from exporting, set HAJER_DISABLED=1 or HAJER_TRACES_ENABLED=0 (see Configuration).
For litellm, attach mode replaces litellm.completion and litellm.acompletion and also rebinds names you imported earlier with from litellm import completion. A google.genai.Client built before attach() runs is not instrumented; pass it to hajer.wrap(client).
Supported libraries
Every row below is generated from the SDK's own target declarations, so the table cannot promise capture the code does not perform. A non-streamed call carries token usage for every library.
- sync / async — a synchronous client or function, and an
asyncone, are both instrumented. - streaming — a streamed call is recorded when the stream is exhausted or closed, never at garbage
collection;
streamChunksandstreamCompletesay what was observed. - usage available — token counts are observable for this library, including for its streamed calls.
A
nois about the streamed path; a non-streamed call carries usage everywhere. - content captured — message text, tool arguments and tool results can be recorded. On by default,
for every library;
HAJER_CAPTURE_CONTENT=0turns it off and keeps shapes, names, counts, timing and usage. - caveat — what to know before trusting the row.
| SDK | sync | async | streaming | usage available | content captured | caveat |
|---|---|---|---|---|---|---|
| openai | yes | yes | yes | no | yes | A streamed call names its tokens only when the request carried stream_options={"include_usage": True}; without it the record says so rather than reporting zero. chat.completions, responses and both stream helpers are covered. |
| anthropic | yes | yes | yes | yes | yes | messages.create, beta.messages.create and messages.stream are covered. A caller who reads only text_stream records 0 chunks; get_final_message() carries the usage. |
| langchain-anthropic | yes | yes | yes | yes | yes | Instrumented at the inner anthropic client the chat model holds (_client / _async_client), so the record is of the request that left the process. |
| langchain-openai | yes | yes | yes | no | yes | Instrumented at the inner openai clients (root_client / root_async_client), through the with_raw_response call LangChain makes: the record is the request that left the process and the answer it parsed. The streamed-usage caveat of openai applies (stream_usage=True). |
| litellm | yes | yes | yes | yes | yes | The upstream provider is read from the model prefix (anthropic/claude-sonnet-5); an unprefixed model is recorded as provider litellm. Patched on the module, so a name bound by from litellm import completion before attach() is rebound — see the import-order rule. |
| google-genai | yes | yes | yes | yes | yes | generate_content_stream is a separate method and is recorded as a stream; usage arrives in the final chunk's usage_metadata. A client constructed before attach() is reached with hajer.wrap(client). |
The SDK never imports a provider library. It finds the methods to instrument by attribute on the object you give it, so installing hajer does not add openai or anthropic to your dependencies. The hajer[openai] and hajer[anthropic] extras exist only for convenience.
Streaming
A streamed call is recorded, and its span emitted, when the stream ends:
- when you exhaust the iterator,
- when you close it, or
- when the
withblock around it exits.
It is never recorded at garbage collection. A stream you never consume, close or wrap in with produces no span.
What you get back is a proxy that forwards the provider's own chunks and methods (text_stream, response, get_final_message(), until_done(), close() and so on). The chunks are the provider's own objects; the stream container is not, so an isinstance check against the provider's stream class fails on it. for loops, next(), an early break and close() behave as before, and exceptions are re-raised unchanged. A stream abandoned partway is recorded with hajer.stream.complete=false. The span keeps the session, user and workflow that were active when the call was opened, wherever you consume the stream.
If your code reads only Anthropic's text_stream, the chunks never pass through the proxy, so hajer.stream.chunks is 0. The call, its model, its timing and its completion are still recorded, and the usage arrives through get_final_message().
For OpenAI, request usage on the final chunk:
stream = client.chat.completions.create(
model="gpt-5",
messages=[{"role": "user", "content": "Write a haiku"}],
stream=True,
stream_options={"include_usage": True}, # without this, token usage is not observable
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")
Without include_usage, the span does not report zero tokens; it records the gap in hajer.limitations.
Raw responses
Calls through with_raw_response and with_streaming_response return the provider's own object, untouched in type and behaviour. The call is recorded when your code reads the body, through any accessor (parse(), read(), json(), text, iter_lines(), iter_bytes(), http_response), and exactly once however many times you read it.
When the SDK cannot read the answer, hajer.limitations says why:
| Limitation | Meaning |
|---|---|
RAW_RESPONSE_NOT_READ | The response was closed before anything read its body, so the answer was not observed. |
RAW_RESPONSE_UNREADABLE | The provider answered 2xx with a body that is not a JSON answer (a proxy's error page, plain text), so no output, usage or finish reason was read. |
RAW_RESPONSE_INCOMPLETE | Your code stopped reading before the end of the body, or the body was larger than HAJER_BODY_MAX_BYTES (default 65536). The record holds what had arrived. |
HAJER_BODY_MAX_BYTES bounds only what the SDK buffers to read the answer. Your code always receives the whole body.
Model requests through your own httpx client
If you call a model API with your own httpx client instead of a provider SDK, wrap its transport:
import httpx
import hajer
http = httpx.Client(transport=hajer.CaptureTransport(httpx.HTTPTransport()))
async_http = httpx.AsyncClient(transport=hajer.AsyncCaptureTransport(httpx.AsyncHTTPTransport()))
A JSON POST to a path ending in /chat/completions, /responses or /messages is recorded as a model call, read from the bytes. Bytes are observed as your code consumes them, never read ahead. The response bytes the SDK retains are bounded by HAJER_WRAPPED_CALL_MAX_BYTES. Headers are never recorded. A request that a wrapped provider client is already making is not recorded twice, and the SDK's own requests are never recorded. The same works for httpx2 (the httpx fork newer openai and anthropic releases send through) when it is installed.
To also record your application's other outbound HTTP requests, see HAJER_CAPTURE_HTTP in Content capture and redaction.
Coexisting with other OpenTelemetry instrumentation
If another OpenTelemetry GenAI instrumentation (for example OpenTelemetry's own openai instrumentor) already traces a call, the SDK does not emit its own span for that call. One call, one span: theirs. The SDK detects this when a span carrying gen_ai.operation.name is current when the call opens, or starts inside the call.
The session, user, tags, metadata, workflow and environment are still added to that instrumentation's spans, each only if the span does not already carry it. See Sessions and context.
Detection is best effort. An instrumentation that sets gen_ai.operation.name only after its span has ended, or opens its span in another context, is not detected. To guarantee no duplicates, turn the SDK's model spans off:
export HAJER_MODEL_SPANS=0
With HAJER_MODEL_SPANS=0, hajer.workflow, hajer.component and hajer.tool spans are still emitted.
What a model span carries
A model span is named {operation} {model}, for example chat gpt-5, embeddings text-embedding-3-small or generate_content gemini-2.5-pro (just {operation} when no model is known). It has kind CLIENT. The span is emitted when the call settles, backdated to when the call began, and is a child of whatever span was current then, typically the workflow or component. It is never made the current span itself, so spans other code opens during the call are not its children.
| Attribute | Meaning |
|---|---|
gen_ai.operation.name | chat (OpenAI chat completions and Responses, Anthropic, litellm), embeddings, or generate_content (google-genai) |
gen_ai.provider.name | openai, anthropic, gcp.gen_ai, the litellm upstream from the model prefix (anthropic/... is anthropic), else litellm |
hajer.api | The surface called: chat.completions, responses, messages, embeddings, litellm.completion, genai.generate_content or genai.generate_content_stream |
gen_ai.request.model, gen_ai.response.model | The model requested and the model that answered |
gen_ai.request.temperature, top_p, top_k, max_tokens, seed, frequency_penalty, presence_penalty, choice.count, stop_sequences | Request settings, when you passed them. max_tokens is read from max_tokens, max_output_tokens or max_completion_tokens; choice.count from n; stop_sequences from stop or stop_sequences. |
hajer.request.tool_names | Names of the tools declared on the request |
gen_ai.response.id, gen_ai.response.finish_reasons | Response id and finish reasons |
gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, gen_ai.usage.total_tokens | Token usage |
hajer.usage.cache_read_input_tokens, hajer.usage.cache_write_input_tokens | Cache token usage, when the provider reports it |
gen_ai.input.messages, gen_ai.output.messages, gen_ai.system_instructions | Content, with content capture on (redacted) |
gen_ai.embeddings.dimension.count, hajer.embedding.vectors | Embedding shape; never the vectors or the input |
server.address, server.port | The provider endpoint, with call-site capture on |
code.function.name, code.file.path, code.line.number | Where in your code the call was made (code.function.name is module.qualified_name), with call-site capture on |
hajer.stream, hajer.stream.complete, hajer.stream.chunks | A streamed call: whether it ran to the end, and how many chunks passed through |
error.type, http.response.status_code | On failure: the exception class and status; never the message |
hajer.limitations | What the record could not observe, one sentence each, for example a stream that reported no usage or content that was clipped |
The span also carries the context attributes described in Sessions and context.
What the SDK cannot capture
At any setting, the SDK does not see:
- what happens to the output after the call, such as a later transformation or template,
- database state, authorization state, or whether a side effect completed,
- retries inside the provider SDK (
WrappedCall.retriesis always0), - a stream you never consume, never close and never wrap in a
withblock.