How to stop CrewAI retries from firing tools twice

CrewAI task retries re-run tools that already succeeded. This post shows the SQLite claim-before-execute guard that stops duplicate payments and emails.

Title card for the CrewAI task retry idempotency guard post
Claim the idempotency key before the side effect so retries return the stored receipt.

TL;DR: CrewAI task retry re-runs any @tool that already succeeded, so claim a stable idempotency key in SQLite before the side effect and return the stored receipt on retry.

A CrewAI crew that charges cards or sends mail will eventually fire the same tool twice. The task fails after the tool succeeds but before the agent records confirmation, the retry loop kicks in, and the tool has no memory of the first run. This bites payment, email, and trade tools first, and most threads about it stop at "set max retries to zero," which trades reliability for silence. The catalog entry for this failure mode sits in the Automation Error Index alongside other agent retry faults.

Why does CrewAI re-run tools on task retry?

CrewAI retries a task when the agent hits an error, when a guardrail rejects the output, or when an external trigger re-fires the run. The default max_retry_limit is 2 and guardrail_max_retries defaults to 3, per the tasks documentation. Each retry replays the task from the agent loop, and the loop has no completion record for individual tool calls. A tool that returned "payment sent" a second ago looks brand new to the next attempt.

The repro in issue #5802 (filed May 2026, still open as of September 2026) is exact: a send_payment tool runs stripe.charge(amount, recipient), the call succeeds, then a failure lands between execution and confirmation. The retry fires the charge again. The issue has drawn over a hundred comments because the shape is familiar to anyone who has run role-based crews in production, as described in the best AI agent framework for Python comparison. The same thread confirms the fix has to live outside the agent process, since in-memory dedup dies with the worker that wrote it.

How to guard a side-effecting tool against duplicate execution?

The guard is a claim-before-execute wrapper with three properties: the key is stable across retries, the claim is written before the side effect, and the store survives process restarts. Do it in this order.

Two-lane flow: without a guard the first attempt fires the charge, a crash lands before confirmation, and the retry fires again for two charges; with a SQLite guard the first attempt claims PENDING, fires once, and commits the receipt, so the retry returns the receipt for one charge
Claim the idempotency key before the side effect runs, so a retry returns the stored receipt instead of firing the tool again.
  1. Derive a semantic key, not a hash of raw args. A payment keys on invoice or order scope plus recipient and rounded amount, such as payment:inv-8841:acme:120.00. An email keys on recipient plus subject plus campaign id. Raw-argument hashing collapses two intentional same-amount payments into one, which is the opposite bug. The tool author owns the key shape because only the tool knows which arguments define a unique logical operation.
  2. Execute, then flip PENDING to COMMITTED with the receipt. Store the provider receipt (charge id, message id, trade ticket) as the receipt value. The receipt is what the retry returns, so it must be the real provider response, not a placeholder string. If the tool raises, leave the row as PENDING for the sweeper or mark it FAILED so the next retry executes cleanly rather than returning a null result.
  3. Wire the PR #5822 opt-in where it fits. The open PR #5822 adds idempotent=True on BaseTool with a pre-claim written before execution and a SQLiteCacheBackend that survives cross-process retries. Mark read-heavy tools idempotent through the flag, and keep the manual SQLite wrapper above for money-moving tools until the PR merges, because the PR default store is still in-memory and the review thread shows the ToolUsage path still uses read-then-add in places.
  4. Put approval or caps in front of irreversible tools. While the framework gap stands, route payments above a threshold through human-in-the-loop confirmation and track per-run spend. The CrewAI hierarchical delegation path has its own sharp edges, such as the DelegateWorkToolSchema validation error, so a second pair of eyes on mutating tools pays twice.

Claim PENDING in SQLite before the side effect runs. Use an atomic insert-if-absent against a table keyed on the idempotency key with columns for status and receipt. If the insert finds an existing COMMITTED row, return the stored receipt without touching the provider. If it finds PENDING from a live worker, wait or raise instead of executing. The minimal schema is (key TEXT PRIMARY KEY, status TEXT, receipt TEXT, updated_at TEXT).

import sqlite3, json, time

def claim(conn, key, timeout_s=300):
    now = time.time()
    try:
        conn.execute(
            "INSERT INTO idempotency(key, status, receipt, updated_at) VALUES (?, 'PENDING', NULL, ?)",
            (key, now),
        )
        conn.commit()
        return "EXECUTE"
    except sqlite3.IntegrityError:
        row = conn.execute(
            "SELECT status, receipt, updated_at FROM idempotency WHERE key = ?", (key,)
        ).fetchone()
        status, receipt, updated = row
        if status == "COMMITTED":
            return ("RETURN", receipt)
        if status == "PENDING" and now - updated < timeout_s:
            return ("WAIT", None)
        conn.execute(
            "UPDATE idempotency SET status='PENDING', updated_at=? WHERE key=?", (now, key)
        )
        conn.commit()
        return "EXECUTE"
Key-shape comparison over two 120-unit payments to the same recipient for different invoices: an amount-only key produces one shared key and the second payment wrongly returns the first receipt, while semantic keys with invoice scope produce two distinct keys so both payments execute exactly once
Key the claim on invoice or order scope plus recipient and rounded amount, never on amount alone, or distinct orders collapse into one.

How to verify the guard worked?

Run the crew twice against the same logical key and count provider-side effects, not log lines. First, trigger a retry on purpose: let the tool succeed, raise once after the commit point, and confirm the retry returns the stored receipt with exactly one charge, one message, or one ticket at the provider. Second, kill the worker between claim and commit and restart on a fresh process pointing at the same SQLite file; the new worker must see the PENDING row and refuse to double-fire. Third, submit two intentionally different operations with near-identical args (two invoices, same amount) and confirm both execute, which proves the key carries scope rather than raw args.

How to recover when the duplicate still happens?

Three causes cover nearly every repeat report. The store is process-local: a dict, LRU cache, or JSONL file next to the worker vanishes on re-dispatch, so move the claim to SQLite on disk or Postgres when workers span hosts. The key misses scope: amount-only keys merge distinct orders, so add invoice, campaign, or trade-window fields. The retry path bypasses the wrapper: guardrail retries and external re-triggers enter through different code paths than max_retry_limit, so apply the decorator at the tool function itself rather than in task callbacks. For deployment shape, the FastAPI and Docker production layout shows where a shared SQLite volume or Postgres sidecar belongs.

FAQ

Does setting max_retry_limit to 0 fix duplicate tool execution?

It hides the symptom by disabling retries, but the first transient error then fails the whole task. Keep the default retry budget and add the idempotency claim, so retries stay safe instead of absent.

Do guardrail retries also re-run tools?

Yes. A guardrail rejection sends the error back to the agent and the task retries up to guardrail_max_retries times, which replays tools the same way an exception retry does. The guard belongs on the tool, not on the retry knob, because every retry source funnels through the same tool function.

What key shape should a payment tool use?

Use payment:{invoice_id}:{recipient}:{rounded_amount}:{currency} with the amount rounded by the tool author (for example to two decimals). Never key on raw floats or on amount alone; rounding and scope are semantic decisions the framework cannot make for you.

Why does the in-memory cache option not survive production retries?

The cache lives inside the worker process, so any retry that lands on a new worker starts with an empty cache and re-fires the tool. The idempotent-tools package documents the same boundary: local SQLite is the minimum durable store, Redis or Postgres when workers are distributed. Claims must be atomic insert-if-absent, not write-after-success.

How do you recover a PENDING claim left by a crashed worker?

Give every PENDING row a timestamp and a sweeper that resets rows older than a timeout (five minutes is a sane default) back to retryable. A crash mid-call should leave PENDING, never a fake COMMITTED; rerunning the side effect once after a timeout beats silently returning nothing.

Does the payment provider's own idempotency key make the wrapper redundant?

No, it complements it. Pass the same logical key down as the provider key (Stripe, for example, accepts one per request), but keep the runtime claim as the outer boundary. The provider key stops double charges; the runtime claim stops double tool bodies, lets the agent return the prior receipt, and covers email and trade tools whose providers have no native key.