How to stop CrewAI retries from firing tools twice
CrewAI task retries re-run tools that already succeeded. This post shows the SQLite claim-before-execute guard that stops duplicate payments and emails.
TL;DR: CrewAI task retry re-runs any @tool that already succeeded, so claim a stable idempotency key in SQLite before the side effect and return the stored receipt on retry.
A CrewAI crew that charges cards or sends mail will eventually fire the same tool twice. The task fails after the tool succeeds but before the agent records confirmation, the retry loop kicks in, and the tool has no memory of the first run. This bites payment, email, and trade tools first, and most threads about it stop at "set max retries to zero," which trades reliability for silence. The catalog entry for this failure mode sits in the Automation Error Index alongside other agent retry faults.
Why does CrewAI re-run tools on task retry?
CrewAI retries a task when the agent hits an error, when a guardrail rejects the output, or when an external trigger re-fires the run. The default max_retry_limit is 2 and guardrail_max_retries defaults to 3, per the tasks documentation. Each retry replays the task from the agent loop, and the loop has no completion record for individual tool calls. A tool that returned "payment sent" a second ago looks brand new to the next attempt.
The repro in issue #5802 (filed May 2026, still open as of September 2026) is exact: a send_payment tool runs stripe.charge(amount, recipient), the call succeeds, then a failure lands between execution and confirmation. The retry fires the charge again. The issue has drawn over a hundred comments because the shape is familiar to anyone who has run role-based crews in production, as described in the best AI agent framework for Python comparison. The same thread confirms the fix has to live outside the agent process, since in-memory dedup dies with the worker that wrote it.
How to guard a side-effecting tool against duplicate execution?
The guard is a claim-before-execute wrapper with three properties: the key is stable across retries, the claim is written before the side effect, and the store survives process restarts. Do it in this order.

- Derive a semantic key, not a hash of raw args. A payment keys on invoice or order scope plus recipient and rounded amount, such as
payment:inv-8841:acme:120.00. An email keys on recipient plus subject plus campaign id. Raw-argument hashing collapses two intentional same-amount payments into one, which is the opposite bug. The tool author owns the key shape because only the tool knows which arguments define a unique logical operation. - Execute, then flip PENDING to COMMITTED with the receipt. Store the provider receipt (charge id, message id, trade ticket) as the receipt value. The receipt is what the retry returns, so it must be the real provider response, not a placeholder string. If the tool raises, leave the row as
PENDINGfor the sweeper or mark itFAILEDso the next retry executes cleanly rather than returning a null result. - Wire the PR #5822 opt-in where it fits. The open PR #5822 adds
idempotent=TrueonBaseToolwith a pre-claim written before execution and aSQLiteCacheBackendthat survives cross-process retries. Mark read-heavy tools idempotent through the flag, and keep the manual SQLite wrapper above for money-moving tools until the PR merges, because the PR default store is still in-memory and the review thread shows theToolUsagepath still uses read-then-add in places. - Put approval or caps in front of irreversible tools. While the framework gap stands, route payments above a threshold through human-in-the-loop confirmation and track per-run spend. The CrewAI hierarchical delegation path has its own sharp edges, such as the DelegateWorkToolSchema validation error, so a second pair of eyes on mutating tools pays twice.
Claim PENDING in SQLite before the side effect runs. Use an atomic insert-if-absent against a table keyed on the idempotency key with columns for status and receipt. If the insert finds an existing COMMITTED row, return the stored receipt without touching the provider. If it finds PENDING from a live worker, wait or raise instead of executing. The minimal schema is (key TEXT PRIMARY KEY, status TEXT, receipt TEXT, updated_at TEXT).
import sqlite3, json, time
def claim(conn, key, timeout_s=300):
now = time.time()
try:
conn.execute(
"INSERT INTO idempotency(key, status, receipt, updated_at) VALUES (?, 'PENDING', NULL, ?)",
(key, now),
)
conn.commit()
return "EXECUTE"
except sqlite3.IntegrityError:
row = conn.execute(
"SELECT status, receipt, updated_at FROM idempotency WHERE key = ?", (key,)
).fetchone()
status, receipt, updated = row
if status == "COMMITTED":
return ("RETURN", receipt)
if status == "PENDING" and now - updated < timeout_s:
return ("WAIT", None)
conn.execute(
"UPDATE idempotency SET status='PENDING', updated_at=? WHERE key=?", (now, key)
)
conn.commit()
return "EXECUTE"
How to verify the guard worked?
Run the crew twice against the same logical key and count provider-side effects, not log lines. First, trigger a retry on purpose: let the tool succeed, raise once after the commit point, and confirm the retry returns the stored receipt with exactly one charge, one message, or one ticket at the provider. Second, kill the worker between claim and commit and restart on a fresh process pointing at the same SQLite file; the new worker must see the PENDING row and refuse to double-fire. Third, submit two intentionally different operations with near-identical args (two invoices, same amount) and confirm both execute, which proves the key carries scope rather than raw args.
How to recover when the duplicate still happens?
Three causes cover nearly every repeat report. The store is process-local: a dict, LRU cache, or JSONL file next to the worker vanishes on re-dispatch, so move the claim to SQLite on disk or Postgres when workers span hosts. The key misses scope: amount-only keys merge distinct orders, so add invoice, campaign, or trade-window fields. The retry path bypasses the wrapper: guardrail retries and external re-triggers enter through different code paths than max_retry_limit, so apply the decorator at the tool function itself rather than in task callbacks. For deployment shape, the FastAPI and Docker production layout shows where a shared SQLite volume or Postgres sidecar belongs.
FAQ
Does setting max_retry_limit to 0 fix duplicate tool execution?
It hides the symptom by disabling retries, but the first transient error then fails the whole task. Keep the default retry budget and add the idempotency claim, so retries stay safe instead of absent.
Do guardrail retries also re-run tools?
Yes. A guardrail rejection sends the error back to the agent and the task retries up to guardrail_max_retries times, which replays tools the same way an exception retry does. The guard belongs on the tool, not on the retry knob, because every retry source funnels through the same tool function.
What key shape should a payment tool use?
Use payment:{invoice_id}:{recipient}:{rounded_amount}:{currency} with the amount rounded by the tool author (for example to two decimals). Never key on raw floats or on amount alone; rounding and scope are semantic decisions the framework cannot make for you.
Why does the in-memory cache option not survive production retries?
The cache lives inside the worker process, so any retry that lands on a new worker starts with an empty cache and re-fires the tool. The idempotent-tools package documents the same boundary: local SQLite is the minimum durable store, Redis or Postgres when workers are distributed. Claims must be atomic insert-if-absent, not write-after-success.
How do you recover a PENDING claim left by a crashed worker?
Give every PENDING row a timestamp and a sweeper that resets rows older than a timeout (five minutes is a sane default) back to retryable. A crash mid-call should leave PENDING, never a fake COMMITTED; rerunning the side effect once after a timeout beats silently returning nothing.
Does the payment provider's own idempotency key make the wrapper redundant?
No, it complements it. Pass the same logical key down as the provider key (Stripe, for example, accepts one per request), but keep the runtime claim as the outer boundary. The provider key stops double charges; the runtime claim stops double tool bodies, lets the agent return the prior receipt, and covers email and trade tools whose providers have no native key.