Write Serialization: Making Parallel Agent Tool Calls Safe to Run Together

Learn how to run agent tool calls in parallel without lost updates: a read/write partition declared in tool schemas, per-resource writer queues that preserve model issue order, reads that wait behind pending writes, conditional writes at the store as the real correctness boundary, bounded waits that keep one hung call from stalling the turn, and order-asserting tests you can run in CI.

Write Serialization: Making Parallel Agent Tool Calls Safe to Run Together

A model turn comes back with three tool calls: fetch the invoice, update its shipping note, charge the card. Modern harnesses fan all three out at once — parallel tool execution is the point, and on a turn with five independent lookups it turns ten seconds of round trips into two. But two of these calls mutate invoice 8123. The update and the charge both read the invoice at revision 4, both write revision 5, and whichever commits last silently erases the other. The customer's card gets charged against a note that no longer exists. Nothing crashed. Every call returned 200.

This is the bug idempotency keys cannot fix. Keys make one logical operation safe to repeat; they say nothing about two different operations running at the same time. Parallel execution adds a second axis of duplication — not twice in sequence, but interleaved — and the remedy is a different mechanism: partition the tool catalog into reads and writes, and make writes to the same resource queue behind each other. Write serialization. This article walks through the classification, the executor-side gate, the store-side boundary that actually guarantees correctness, and the ways a naive gate stalls or deadlocks.

What actually races

Not everything in a parallel turn races. Sort every pair of calls by what they touch:

  • Read after read never races. Concurrent lookups are the legitimate win of parallel turns.
  • Write to disjoint resources is safe — assign_ticket(t_9) and charge_card(inv_4) cannot lose data to each other.
  • Write after write on the same resource is a lost update. Two read-modify-write cycles interleave; the second commits against state the first already replaced.
  • Read after write on the same resource is a stale read — and it is worse when the read feeds a later write, because the agent then makes a decision on data that was already wrong when it arrived.

Agents make these hazards more likely, not less, for a structural reason: the model has no idea your executor is concurrent. A model emits a turn as an ordered list of intentions; it writes update_invoice before charge_card because the note should exist before the money moves. The harness treats that list as a set and fans it out. The moment it does, the model's ordering intent is discarded — and the model cannot restore it, because nothing in the transcript looks wrong. The harness owns execution ordering exactly the way it owns retries: silently, structurally, and with nobody else to blame.

Declare the partition in the tool schema

The gate needs two facts per tool: does it mutate, and which argument identifies the mutated resource. Both belong in the tool schema, declared once by the tool's author — not inferred at runtime by the executor, and not guessed from the tool's name. Tool schemas are contracts the model reads and the executor enforces; this is the executor's half.

{
  "name": "update_invoice",
  "description": "Update mutable fields of an existing invoice.",
  "parameters": {
    "invoice_id": "string",
    "shipping_note": "string"
  },
  "x-mutates": true,
  "x-resource-key": ["invoice", "$args.invoice_id"]
}

The x- prefix keeps the declaration out of the parameters the model sees; the model should not have to think about concurrency, and with this contract it never does. Reads declare nothing and stay concurrent. A tool that both reads and writes is declared as a write — conservative, and correct.

The resource key deserves the same care as an idempotency key's scope. invoice:8123 is the key; "8123" alone is not, because two different tools may have colliding ids. Derive the key from the declared expression — never from the model, never from a hash of the whole payload (two updates to the same invoice with different notes must serialize, so the key cannot depend on the note).

One writer per key, in issue order

With the partition declared, the executor becomes small. Reads on untouched keys run immediately. Everything else joins a per-key FIFO queue that preserves the order the model issued:

const tails = new Map<string, Promise<unknown>>();

function enqueue(call: ToolCall, run: () => Promise<unknown>): Promise<unknown> {
  const meta = registry[call.name];          // from the tool schemas
  const key = meta.keyOf(call.args);         // e.g. "invoice:8123"
  const prev = tails.get(key);

  // A read on a key with no pending work runs concurrently.
  // Anything touching a busy key joins that key's queue, in issue order.
  if (!prev && !meta.mutates) return run();

  const next = (prev ?? Promise.resolve())
    .catch(() => {})   // a failed predecessor must not poison the chain
    .then(run);
  tails.set(key, next);
  return next;
}

async function executeTurn(calls: ToolCall[]) {
  const settled = await Promise.all(
    calls.map(async (c) => [c.id, await enqueue(c, () => invoke(c))]),
  );
  return Object.fromEntries(settled);
}

Three details carry the correctness. Issue order comes free: the turn's array order is the model's intended order, and FIFO per key preserves it. The .catch matters because a chain whose head rejected would skip everything queued behind it — one failing write must not cancel the writes the model issued after it. And reads queue behind writes on their key (!prev fails, so the read chains) because a read that skips the queue observes pre-write state — exactly the stale-read hazard. Consecutive reads on the same key also serialize here; that is unnecessary but harmless, and you can batch them into a shared wave if profiling ever cares.

What you must not do is serialize globally. One process-wide lock for all writes is correct and destroys the latency win that justified parallelism in the first place — five independent mutations become five round trips in a row. Per-key queues keep independent work parallel and only serialize the calls that were actually racing.

The gate is an optimization; the store is the boundary

An in-process queue covers exactly one process. Retries redeliver tool calls, agent replicas run behind the same load balancer, and a crash mid-write re-runs the turn on another worker. The correctness boundary has to live where the data lives, with conditional writes:

UPDATE invoices
SET    revision = revision + 1,
       shipping_note = $note
WHERE  invoice_id = $id
  AND  revision = $seen_revision;
-- zero rows affected: someone wrote first. Reload, re-decide, retry loudly.

This is optimistic concurrency control — the earlier deep-dive on version columns, ETags, and compare-and-swap covers the store side in full. What the agent layer adds is the right failure behavior: a lost-update conflict surfaces in the transcript as a 409, the model re-reads the invoice, sees the note it raced against, and decides again. A silent last-writer-wins turns the same race into corrupted state that no one — model or human — ever gets to reconsider. Loud beats lost.

The gate and the store do different jobs. The store makes concurrent writes safe; the gate makes them rare and ordered, so the model's issue order survives instead of every conflict paying a re-read round trip in the common case. Deploy the store first; the gate is the performance half.

Don't hold a key across the world

Two traps turn a correct gate into a stalled one.

The hung write. A tool call that hangs holds its key, and every call queued behind it inherits the hang. Bound every write with a timeout derived from the run's remaining deadline — the deadline-propagation discipline from the resource-budget piece — so a stuck call fails after its slice and the queue advances without it. The hung call itself should be cancelled if the transport allows it; if it eventually lands anyway, the store's version check is what catches it. This is another reason the gate cannot be the only mechanism.

The approval gate inside a write. A mutation that pauses for human approval holds its key for the length of a lunch break — or a weekend. Approval gates belong between tool calls, not inside them. Split the mutation into a reserve step (validate, lock the version, mark pending) and a confirm step (apply after approval), so the key is held for milliseconds around each. The ledger pattern from earlier in this series is the natural backbone for the pending state.

What write serialization does not fix

  • Cross-resource business ordering. "Charge the card, then ship the order" spans two keys; no per-key queue relates them. Enforcing that sequence is workflow territory — sagas and compensating actions, covered earlier in this series.
  • Wrong order issued by the model. The gate preserves issue order; it does not choose it. A model that updates after charging gets a faithfully preserved wrong order, which is at least visible in the transcript.
  • Turn-boundary hazards. Sequential turns are already ordered by the conversation loop. The race only exists inside a turn — which is also why the fix can stay there.

Testing: assert order, not just outcomes

Final-state assertions pass even when ordering is broken, if the writes happen to interleave benignly. Assert the execution order per key, and inject hangs to prove the gate degrades instead of stalling:

async def test_parallel_turn_preserves_issue_order(gateway):
    await gateway.execute_turn(turn(
        ("c1", "update_invoice", {"invoice_id": "8123", "note": "desk"}),
        ("c2", "charge_card", {"invoice_id": "8123", "amount": 4900}),
    ))
    assert gateway.order_log("invoice:8123") == ["update_invoice", "charge_card"]
    assert invoice("8123").note == "desk"      # lost update would erase this

async def test_hung_write_does_not_stall_the_turn(gateway, hang):
    hang("update_invoice", seconds=30)
    result = await asyncio.wait_for(
        gateway.execute_turn(turn(
            ("c1", "update_invoice", {"invoice_id": "8123", "note": "desk"}),
            ("c2", "charge_card", {"invoice_id": "8123", "amount": 4900}),
        )),
        timeout=5,
    )
    assert result["c1"].status == "timeout"    # bounded, and reported
    assert result["c2"].status == "ok"         # queue advanced past it

A third test belongs at the store: two conditional writes against the same revision, one must fail. That is the boundary the whole pattern leans on, and it is two lines with a version column.

The checklist

  • Tool schemas declare mutating tools and a resource-key expression; reads stay concurrent, writes are keyed.
  • Per-key FIFO queues preserve the model's issue order; a failed head never poisons the chain.
  • Reads on busy keys wait behind pending writes; independent keys never wait on each other.
  • The store enforces conditional writes — the gate orders, the store guarantees.
  • Every write carries a timeout derived from the run deadline, and no key is ever held across an approval gate.
  • CI asserts per-key execution order and lost-update absence under injected hangs — not just final state.

The previous chapter made repeats safe; this one makes neighbors safe. An agent whose harness fans out independent calls and still mutates each resource like a single careful thread keeps both the latency and the correctness — and neither the model nor the tool author has to think about it again.

About this story

This article was written by Gen-AI using GPT 5.5 or Opus 4.7. Verify technical guidance before using it in production systems.

Advertisement ad.endcap · 336×280