LangChain ReAct Agent Token and Cost Calculator
Multi-step agent runs are billed per step, and retrieved context is paid for on every one of them. Set the four inputs below to see tokens per execution and cost per 1,000 runs.
How this estimate is calculated
Input tokens per run are steps × (base prompt + retrieved context). Output tokens
are assumed to be 100 per step, a typical length for a short reasoning turn plus a tool call.
The two figures are priced separately, because input and output are billed at different rates.
Model rates are the published list prices shown in the selector and are not fetched live: GPT-4o
at $2.50 in / $10.00 out per 1M tokens per the
OpenAI API pricing page,
and Claude Haiku 4.5 at $1 in / $5 out per 1M tokens per the
Anthropic pricing page.
Both exclude prompt caching and batch discounts, which cut input cost substantially on a workload
that resends the same system prompt every step. Check the provider's current page before budgeting
against these numbers.
The estimate holds the prompt flat across steps. A real ReAct agent appends every prior action and observation to its scratchpad, so later steps carry more tokens than earlier ones. Treat the result as a floor rather than a ceiling, and treat a high step count as the variable worth attacking first: it multiplies against the entire prompt, retrieved context included.
Retrieval share is the fraction of input tokens taken up by retrieved context. When it climbs past roughly half, the cheapest available saving is usually retrieving fewer and better-ranked chunks rather than switching to a cheaper model.
Related guides
- LangChain RAG pipeline: setup to first answer: choosing chunk size and retrieval depth, the two inputs that set the context figure above.
- LangChain building blocks: why a deterministic chain often replaces an agent loop, and removes most of the step count with it.
- LangChain agent errors: capping runaway loops before they multiply the step count without bound.
- LangChain vs LangGraph: agents, state and control flow: how the orchestration choice changes the number of model calls a task needs.