· AI Engineering · 12 min read
Agent Reliability Patterns: Retries, Idempotency and Observability in Agentic Loops
Agent reliability patterns fail when retries re-run side effects. Design idempotency keys and step-level tracing into the loop from day one.

In my work with teams moving agents from demo to production, the first serious incident is rarely a bad answer. It is a duplicate. A refund issued twice, a ticket opened three times, an email sent to a customer on every retry of a step that “failed” after the message had already gone out. The model did nothing exotic. The platform did what platforms have always done when a call times out: it tried again.
Retries are the oldest reliability tool we have, and they work because we assume the retried operation is the same operation. In a classic service, you resend the same request with the same payload and get the same effect, or you design it so you do. In an agent loop that assumption quietly fails. The step being retried is a model call that may choose a different action the second time, and the actions it chooses often touch the outside world.
This article is about what breaks, and what to build so it does not. The short version: idempotency and tracing are not features you bolt on after the first incident. They are properties of the loop’s design, and they decide whether an agent is something you can put in front of customers or something you babysit.
- A retry in an agent loop re-runs a non-deterministic step, so it can pick a different action, not just repeat the same one.
- Derive idempotency keys from the run and the intended action, not from the model's output text, and enforce them at the tool boundary.
- Separate retrying the thinking (safe to repeat) from retrying the doing (only safe with a key and a recorded result).
- Step-level traces with inputs, outputs, tool arguments and decisions are what let you answer "what did it do, and why" after the fact. Budget for them up front.
Why classic retry logic misleads you here
Take a standard retry policy: exponential backoff, jitter, a cap of three attempts. It is a good policy for a service call. It is a risky one for a loop step, for three reasons.
First, the step is not a pure function. If you resend the same prompt to a model, you may get a different tool call. Even at low temperature, small changes in context, such as an extra line of tool output from the first attempt, can move the model somewhere else. So “retry” does not mean “do the same thing again”. It means “ask again and see what happens”.
Second, a failure is ambiguous. A timeout on a tool call tells you the caller stopped waiting. It does not tell you whether the tool ran. Distributed systems engineers have lived with this for decades, but in an agent loop the ambiguity is worse because the caller is a language model that will happily narrate the failure and try something creative.
Third, the loop carries state forward. By the time a step fails, the conversation already holds earlier tool results, partial plans, and intermediate decisions. A blind retry of the whole run replays all of that. A retry of just the last step has to reconstruct exactly the context the model saw, or it is no longer the same step.
None of this means retries are bad. It means a retry needs to know what kind of step it is retrying.
Two kinds of steps: thinking and doing
The most useful distinction I have found is between steps that only compute and steps that change the world.
Thinking steps are model calls that read context and produce text, a plan, or a proposed tool call. Repeating them costs tokens and time, and may produce a different proposal, but nothing outside the loop has changed. These are safe to retry aggressively, with one caveat: if the first attempt’s output was partially consumed, say a streamed tool call was already dispatched, treat it as a doing step.
Doing steps are tool executions with side effects: writing a record, sending a message, moving money, calling a third-party API that bills per call. These are the ones that need protection, and the protection is not a smarter retry policy. It is a guarantee at the point of execution that doing it twice is the same as doing it once.
Once you draw the line, the architecture follows. The loop can retry thinking freely. Doing is routed through a thin layer that owns idempotency, records results, and refuses to repeat itself.
Idempotency keys: where they come from
An idempotency key tells the receiving system “this is the same request as before, do not apply it again”. The payment world has used this for years, and the pattern transfers directly. The question for agents is what the key is derived from.
The tempting answer is a hash of the tool arguments. It is wrong more often than you would expect. Two legitimate actions can have identical arguments (refund this customer 20 dollars, twice, for two separate orders), and two retries of one intended action can have different arguments if the model rephrased a free-text field on the second attempt.
A better answer ties the key to the intent, not the text:
- the run identifier,
- the step index or plan item the action belongs to,
- the target resource (the order, the ticket, the account),
- and the action name.
Together these say “in this run, for this planned step, perform this action on this resource”. A retry that lands on the same planned step produces the same key, whatever words the model used. A genuinely different action produces a different key.
Here is the shape of it in plain Python, with no library assumptions:
def execute_tool(run_id, step_id, tool, target, args, store, backend):
key = f"{run_id}:{step_id}:{tool.name}:{target}"
prior = store.get(key)
if prior is not None:
# Already done (or in flight): return the recorded outcome.
return prior.result
store.put_pending(key) # atomic; fails if another worker won
try:
result = backend.call(tool, target, args, idempotency_key=key)
except Exception as err:
store.put_unknown(key, err) # do not assume it did not run
raise
store.put_done(key, result)
return resultThree details matter here. The pending record must be written atomically before the call, so two workers cannot both proceed. A failure is recorded as unknown, not as “not done”, because a timeout does not prove the tool did nothing. And wherever the downstream system supports its own idempotency key, pass yours through, since your store only protects you from your own retries and not from the network between you and the vendor.
Handling the unknown outcome
That “unknown” state is where most agent incidents live. The honest options are limited, and you should pick one deliberately per tool rather than let the model improvise.
- Reconcile. Query the downstream system for the effect before doing anything else. Did the ticket get created? Is the refund on the ledger? This works when the tool has a read counterpart, and it is the best option when it exists.
- Replay with the same key. If the downstream system honors idempotency keys, resend the identical request. The vendor deduplicates.
- Escalate. If neither works, stop the run and hand it to a person with the full trace. For high-consequence actions, this is not a failure of the design. It is the design.
What you should not do is let the model see “tool call timed out” and decide for itself. A model asked what to do after an ambiguous failure will often choose to try again, which is exactly the wrong move for a non-idempotent tool. Give it a structured status instead, something like “outcome unknown, reconciliation pending”, and let the executor own the next move.
Seeing the loop
The diagram below shows where each concern sits. Thinking is retried inside the loop. Doing passes through the idempotency gate. Every transition writes a trace event.

Notice what the model never does: it never decides whether a doing step is safe to repeat. That decision lives in code that can be tested.
Observability that answers the right question
When an agent misbehaves, the question is not “did the service return 200”. It is “what did it do, in what order, based on what it saw, and why did it choose that”. Ordinary request logs cannot answer that. Step-level tracing can.
A useful trace for each step records:
- the run and step identifiers, plus the attempt number, so you can see retries as retries,
- the exact context the model saw, or a reference to a stored snapshot of it,
- the model’s output, including the proposed tool call and arguments,
- the idempotency key and what the gate decided (new, duplicate, unknown),
- the tool’s raw result and latency,
- and the model identifier and parameters, so a behavior change after a model update is attributable.
Treat the attempt number as a first-class field. Many of the worst debugging sessions I have sat through came down to the team not knowing a step had run four times, because the logs showed four similar lines with no link between them.
Two cautions. Traces contain prompts and tool results, which often contain customer data, so decide retention and redaction up front, and apply the same access controls you would to the underlying records. And tracing has a cost in storage and in latency if you write synchronously, so sample read-only steps if you must, but keep every doing step, always. That is the one place where thrift is false economy.
A worked example
Consider a support agent that can look up an order, issue a refund, and email the customer. A customer writes in about a damaged item.
The agent reads the order (thinking plus a read-only tool, freely retryable). It decides to refund 40 dollars and calls the refund tool. The key is built from the run, plan step 3, the order identifier, and the action name. The gate records pending and calls the payments API with that key. The call times out after the payment processor has in fact applied the refund.
In a naive loop, the timeout reaches the model, which concludes the refund failed and calls the tool again. Depending on how the arguments are worded, the second call might not even match the first, and the customer is refunded twice.
In the designed loop, the gate records the outcome as unknown. The executor reconciles by querying the payments system for refunds against that order, finds one, marks the key as done with that result, and returns it to the model as a success. The email step then runs under its own key, so if the mail provider times out, the customer does not get three apologies. The trace shows one refund, one reconciliation, and one email, with every attempt visible.
The model was never smarter in the second version. The surrounding system was just less trusting.
What this costs, and where it is overkill
This is not free. You are adding a store, a gate, reconciliation logic per tool, and a tracing pipeline. For a read-only research agent, or an internal assistant that drafts text a human reviews before anything happens, most of it is overkill. A plain retry on the model call is fine, and the cost of a duplicate is nothing.
The investment pays when three things are true: the agent takes actions, those actions are hard or expensive to reverse, and nobody reviews each one before it lands. Money movement, customer communication, infrastructure changes and anything that triggers billable calls all qualify.
There is also a model-side choice that interacts with this. Smaller or self-hosted models may need more retries for malformed tool calls, which makes the thinking-versus-doing split more important, since you will retry more often. Frontier-hosted models may fail less on format but add network hops and provider-side timeouts you do not control, which makes the unknown-outcome path matter. Either way the same gate and the same traces apply. The pattern does not depend on which model sits behind the loop.
The business case in one paragraph
For a leader deciding whether to fund this work, the framing is simple. Agent incidents are expensive in an unusual way: not outages, but trust. A duplicated refund is small money, but a customer who gets three emails, or an auditor who asks “show me what the system did on this account” and gets a shrug, is a larger cost. Idempotency limits the blast radius of any single failure. Tracing turns “we think it was the model” into a specific, fixable finding, and shortens every post-incident review. Both are cheaper to build before launch than to retrofit once real customers have found the edge cases for you.
A short checklist to take away
Before an agent loop goes to production, ask:
- Can every tool be classified as read-only or side-effecting, and is that enforced in code rather than by prompt?
- Does every side-effecting call carry a key derived from run, step, target and action?
- Is there a defined behavior, reconcile, replay or escalate, for each tool when the outcome is unknown?
- Does the model receive structured failure statuses instead of raw errors it can improvise around?
- Can you reconstruct any past run step by step, including retries and the context at each point?
- Are trace retention and redaction decided, not defaulted?
If you can answer yes to all six, retries stop being a gamble. They go back to being what they were meant to be: a cheap way to ride out a bad moment, applied to steps that can safely absorb them.



