Skip to main content

Instrumenting clients

The SDK records a model call at the provider client that makes it. Each recorded call becomes one gen_ai span (a generation on the platform) when the call settles. There are three ways to instrument clients.

ApproachWhat it instrumentsCode changeUse it when
hajer.wrap(client)One client object you pass inOne line where you build the clientYou build your clients in a few known places.
hajer.instrument()Every openai and anthropic client built afterwardsOne line at startupYou build clients in many places, or inside libraries.
Attach mode (hajer.attach())Every client of every supported library built afterwardsNone, or one importYou want telemetry from an application without editing it.

All three record calls the same way, and you only need one of them.

wrap: one client​

import hajer
from openai import AsyncOpenAI, OpenAI
from anthropic import Anthropic

openai_client = hajer.wrap(OpenAI())
async_client = hajer.wrap(AsyncOpenAI())
anthropic_client = hajer.wrap(Anthropic())

wrap instruments the object in place and returns it. Calling it twice on the same client has no effect.

wrap also accepts a google.genai.Client and the litellm module:

import litellm
from google import genai

import hajer

gemini = hajer.wrap(genai.Client())
hajer.wrap(litellm) # patches litellm.completion and litellm.acompletion

If wrap does not recognise the object, it raises hajer.UnsupportedClientError rather than silently recording nothing. A unittest.mock object is the exception: it is returned unchanged and records nothing, so tests that patch the provider class keep working.

LangChain chat models​

Wrap the chat model itself. The SDK instruments the provider client the chat model holds and returns the same chat model, so you can pass it anywhere that accepts a BaseChatModel, including after bind_tools.

from langchain_openai import ChatOpenAI
from langchain_anthropic import ChatAnthropic

import hajer

llm = hajer.wrap(ChatOpenAI(model="gpt-5"))
claude = hajer.wrap(ChatAnthropic(model="claude-sonnet-5"))

reply = llm.invoke("Summarize this ticket")

The span records the request that left your process, not LangChain's internal representation of it. For ChatOpenAI the SDK instruments root_client and root_async_client, including the with_raw_response calls LangChain makes; for ChatAnthropic, _client and _async_client. A chat model that keeps its provider client anywhere else raises UnsupportedClientError rather than recording nothing.

instrument(): every client built from now on​

import hajer

hajer.instrument() # call once, before you build clients

from openai import OpenAI

client = OpenAI() # instrumented at construction

instrument() patches the constructors of these classes so every instance is instrumented when it is built:

  • openai: OpenAI, AsyncOpenAI, AzureOpenAI, AsyncAzureOpenAI
  • anthropic: Anthropic, AsyncAnthropic, AnthropicBedrock, AsyncAnthropicBedrock, AnthropicVertex, AsyncAnthropicVertex

This also covers LangChain chat models, because they build these clients internally. If the library is not imported yet, instrument() installs an import hook and patches it when it is imported.

instrument() returns an Instrumentation receipt. A second call returns the same receipt. hajer.uninstrument() restores the original constructors and removes the import hook.

hajer.instrument()
...
hajer.uninstrument()
note

instrument() does not cover google-genai or litellm. Use hajer.wrap(...) or attach mode for those.

Attach mode: no code change​

Attach mode instruments every library in the support table for the whole process. It is opt-in. There are three ways to turn it on, and they all run the same hook.

From the shell, with no change to your code:

PYTHONPATH="$(python -m hajer attach-path)" HAJER_ATTACH=1 python -m your_app

hajer attach-path prints a directory containing a sitecustomize.py that Python imports at startup. Python imports only the first sitecustomize it finds, so if your environment already has one, this would shadow it. In that case use the one-line import below instead.

With one import in your entry point:

import hajer.autoattach # noqa: F401 attaches only when HAJER_ATTACH is on

This line reads HAJER_ATTACH, so you can leave it in your code and use the variable as the switch.

In code, unconditionally:

import hajer

attachment = hajer.attach()
print(attachment.describe())
# hajer attached to 10 client classes (openai, anthropic); exporting every model call as a span

hajer.attach() does not read HAJER_ATTACH. It is idempotent: a second call returns the standing attachment. It returns an Attachment with:

  • classes: the client classes it patched
  • modules: the libraries it found
  • inert: True when no key is configured, so nothing will be sent
  • describe(): a one-line summary for a log

Libraries that are already imported are patched immediately. Libraries imported later are patched by an import hook.

hajer.attachment() returns the current Attachment, or None when the process is not attached.

hajer.detach() removes the import hook and stops instrumenting clients built afterwards. It does not un-patch anything: clients that are already instrumented keep recording and exporting. To stop a process from exporting, set HAJER_DISABLED=1 or HAJER_TRACES_ENABLED=0 (see Configuration).

note

For litellm, attach mode replaces litellm.completion and litellm.acompletion and also rebinds names you imported earlier with from litellm import completion. A google.genai.Client built before attach() runs is not instrumented; pass it to hajer.wrap(client).

Supported libraries​

Every row below is generated from the SDK's own target declarations, so the table cannot promise capture the code does not perform. A non-streamed call carries token usage for every library.

  • sync / async — a synchronous client or function, and an async one, are both instrumented.
  • streaming — a streamed call is recorded when the stream is exhausted or closed, never at garbage collection; streamChunks and streamComplete say what was observed.
  • usage available — token counts are observable for this library, including for its streamed calls. A no is about the streamed path; a non-streamed call carries usage everywhere.
  • content captured — message text, tool arguments and tool results can be recorded. On by default, for every library; HAJER_CAPTURE_CONTENT=0 turns it off and keeps shapes, names, counts, timing and usage.
  • caveat — what to know before trusting the row.
SDKsyncasyncstreamingusage availablecontent capturedcaveat
openaiyesyesyesnoyesA streamed call names its tokens only when the request carried stream_options={"include_usage": True}; without it the record says so rather than reporting zero. chat.completions, responses and both stream helpers are covered.
anthropicyesyesyesyesyesmessages.create, beta.messages.create and messages.stream are covered. A caller who reads only text_stream records 0 chunks; get_final_message() carries the usage.
langchain-anthropicyesyesyesyesyesInstrumented at the inner anthropic client the chat model holds (_client / _async_client), so the record is of the request that left the process.
langchain-openaiyesyesyesnoyesInstrumented at the inner openai clients (root_client / root_async_client), through the with_raw_response call LangChain makes: the record is the request that left the process and the answer it parsed. The streamed-usage caveat of openai applies (stream_usage=True).
litellmyesyesyesyesyesThe upstream provider is read from the model prefix (anthropic/claude-sonnet-5); an unprefixed model is recorded as provider litellm. Patched on the module, so a name bound by from litellm import completion before attach() is rebound — see the import-order rule.
google-genaiyesyesyesyesyesgenerate_content_stream is a separate method and is recorded as a stream; usage arrives in the final chunk's usage_metadata. A client constructed before attach() is reached with hajer.wrap(client).

The SDK never imports a provider library. It finds the methods to instrument by attribute on the object you give it, so installing hajer does not add openai or anthropic to your dependencies. The hajer[openai] and hajer[anthropic] extras exist only for convenience.

Streaming​

A streamed call is recorded, and its span emitted, when the stream ends:

  • when you exhaust the iterator,
  • when you close it, or
  • when the with block around it exits.

It is never recorded at garbage collection. A stream you never consume, close or wrap in with produces no span.

What you get back is a proxy that forwards the provider's own chunks and methods (text_stream, response, get_final_message(), until_done(), close() and so on). The chunks are the provider's own objects; the stream container is not, so an isinstance check against the provider's stream class fails on it. for loops, next(), an early break and close() behave as before, and exceptions are re-raised unchanged. A stream abandoned partway is recorded with hajer.stream.complete=false. The span keeps the session, user and workflow that were active when the call was opened, wherever you consume the stream.

If your code reads only Anthropic's text_stream, the chunks never pass through the proxy, so hajer.stream.chunks is 0. The call, its model, its timing and its completion are still recorded, and the usage arrives through get_final_message().

For OpenAI, request usage on the final chunk:

stream = client.chat.completions.create(
model="gpt-5",
messages=[{"role": "user", "content": "Write a haiku"}],
stream=True,
stream_options={"include_usage": True}, # without this, token usage is not observable
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")

Without include_usage, the span does not report zero tokens; it records the gap in hajer.limitations.

Raw responses​

Calls through with_raw_response and with_streaming_response return the provider's own object, untouched in type and behaviour. The call is recorded when your code reads the body, through any accessor (parse(), read(), json(), text, iter_lines(), iter_bytes(), http_response), and exactly once however many times you read it.

When the SDK cannot read the answer, hajer.limitations says why:

LimitationMeaning
RAW_RESPONSE_NOT_READThe response was closed before anything read its body, so the answer was not observed.
RAW_RESPONSE_UNREADABLEThe provider answered 2xx with a body that is not a JSON answer (a proxy's error page, plain text), so no output, usage or finish reason was read.
RAW_RESPONSE_INCOMPLETEYour code stopped reading before the end of the body, or the body was larger than HAJER_BODY_MAX_BYTES (default 65536). The record holds what had arrived.

HAJER_BODY_MAX_BYTES bounds only what the SDK buffers to read the answer. Your code always receives the whole body.

Model requests through your own httpx client​

If you call a model API with your own httpx client instead of a provider SDK, wrap its transport:

import httpx
import hajer

http = httpx.Client(transport=hajer.CaptureTransport(httpx.HTTPTransport()))
async_http = httpx.AsyncClient(transport=hajer.AsyncCaptureTransport(httpx.AsyncHTTPTransport()))

A JSON POST to a path ending in /chat/completions, /responses or /messages is recorded as a model call, read from the bytes. Bytes are observed as your code consumes them, never read ahead. The response bytes the SDK retains are bounded by HAJER_WRAPPED_CALL_MAX_BYTES. Headers are never recorded. A request that a wrapped provider client is already making is not recorded twice, and the SDK's own requests are never recorded. The same works for httpx2 (the httpx fork newer openai and anthropic releases send through) when it is installed.

To also record your application's other outbound HTTP requests, see HAJER_CAPTURE_HTTP in Content capture and redaction.

Coexisting with other OpenTelemetry instrumentation​

If another OpenTelemetry GenAI instrumentation (for example OpenTelemetry's own openai instrumentor) already traces a call, the SDK does not emit its own span for that call. One call, one span: theirs. The SDK detects this when a span carrying gen_ai.operation.name is current when the call opens, or starts inside the call.

The session, user, tags, metadata, workflow and environment are still added to that instrumentation's spans, each only if the span does not already carry it. See Sessions and context.

Detection is best effort. An instrumentation that sets gen_ai.operation.name only after its span has ended, or opens its span in another context, is not detected. To guarantee no duplicates, turn the SDK's model spans off:

export HAJER_MODEL_SPANS=0

With HAJER_MODEL_SPANS=0, hajer.workflow, hajer.component and hajer.tool spans are still emitted.

What a model span carries​

A model span is named {operation} {model}, for example chat gpt-5, embeddings text-embedding-3-small or generate_content gemini-2.5-pro (just {operation} when no model is known). It has kind CLIENT. The span is emitted when the call settles, backdated to when the call began, and is a child of whatever span was current then, typically the workflow or component. It is never made the current span itself, so spans other code opens during the call are not its children.

AttributeMeaning
gen_ai.operation.namechat (OpenAI chat completions and Responses, Anthropic, litellm), embeddings, or generate_content (google-genai)
gen_ai.provider.nameopenai, anthropic, gcp.gen_ai, the litellm upstream from the model prefix (anthropic/... is anthropic), else litellm
hajer.apiThe surface called: chat.completions, responses, messages, embeddings, litellm.completion, genai.generate_content or genai.generate_content_stream
gen_ai.request.model, gen_ai.response.modelThe model requested and the model that answered
gen_ai.request.temperature, top_p, top_k, max_tokens, seed, frequency_penalty, presence_penalty, choice.count, stop_sequencesRequest settings, when you passed them. max_tokens is read from max_tokens, max_output_tokens or max_completion_tokens; choice.count from n; stop_sequences from stop or stop_sequences.
hajer.request.tool_namesNames of the tools declared on the request
gen_ai.response.id, gen_ai.response.finish_reasonsResponse id and finish reasons
gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, gen_ai.usage.total_tokensToken usage
hajer.usage.cache_read_input_tokens, hajer.usage.cache_write_input_tokensCache token usage, when the provider reports it
gen_ai.input.messages, gen_ai.output.messages, gen_ai.system_instructionsContent, with content capture on (redacted)
gen_ai.embeddings.dimension.count, hajer.embedding.vectorsEmbedding shape; never the vectors or the input
server.address, server.portThe provider endpoint, with call-site capture on
code.function.name, code.file.path, code.line.numberWhere in your code the call was made (code.function.name is module.qualified_name), with call-site capture on
hajer.stream, hajer.stream.complete, hajer.stream.chunksA streamed call: whether it ran to the end, and how many chunks passed through
error.type, http.response.status_codeOn failure: the exception class and status; never the message
hajer.limitationsWhat the record could not observe, one sentence each, for example a stream that reported no usage or content that was clipped

The span also carries the context attributes described in Sessions and context.

What the SDK cannot capture​

At any setting, the SDK does not see:

  • what happens to the output after the call, such as a later transformation or template,
  • database state, authorization state, or whether a side effect completed,
  • retries inside the provider SDK (WrappedCall.retries is always 0),
  • a stream you never consume, never close and never wrap in a with block.