Home Services Work About Blog FAQ Contact
← Back to blog

The cost of one bad data integration (a postmortem).

Eleven months ago we shipped an inventory sync between a client's store and their fulfillment partner. It worked cleanly for six weeks. Then, over one weekend, it quietly shipped forty-three orders twice.

Nobody caught it until Monday morning, and by then customers had already started emailing asking why two of the same sweater showed up on their porch. Here's what actually broke, what it cost, and what we rebuilt afterward.

What the integration was supposed to do

The workflow was simple on paper. Shopify fires an order-paid webhook. An n8n workflow catches it, formats the line items, and calls the fulfillment partner's API to create a shipment request. When the partner confirms, n8n writes the tracking number back to the Shopify order. Three systems, one job each. We'd built a dozen integrations shaped exactly like this one before, and none of them had ever given us trouble.

Where it actually broke

The partner's API normally answered in under two seconds. On the Saturday of a site-wide sale, their endpoint slowed down — twelve, sometimes fifteen seconds under the load. Our HTTP Request node had a ten-second timeout with three retries built in. Standard setup. The kind of thing you configure once and stop thinking about.

Here's the part we missed. The first call was often succeeding on the partner's side before our node gave up waiting for a response. The shipment got created. Our workflow just never heard back in time, so n8n fired its built-in retry — which generated a fresh request ID for the new attempt and tried again. The partner's own deduplication checked that request ID, not the Shopify order number. A retry with a new ID looked like a brand new order every time it ran.

Three systems, each doing exactly what we told it to do. Nothing crashed. Nothing logged an error. The order just got shipped twice, and every piece of the chain reported success.

How we actually found out

Not from a dashboard. Not from an alert. A customer emailed the client asking why she'd been charged once and received two identical orders, and the client forwarded it to us Monday at 9am with one line: "is this us?" It was.

We pulled the n8n execution logs first, expecting to find a workflow that had crashed and restarted badly. Instead we found forty-three pairs of successful runs, each one logged as a clean completion, timestamps anywhere from eight to forty seconds apart. That gap is what gave it away — too short to be a separate order, too long to be a simple double-click on the storefront. Once we knew what we were looking for, matching it to the partner's slow Saturday was fast. Knowing to look in the first place took a customer complaint we should have caught ourselves.

What it actually cost

Forty-three double-shipped orders over one weekend, averaging around $38 in product plus shipping on the duplicate. Call it just under $1,900 once the refunds and credits went out. That's the easy number to put on an invoice.

The harder one is the six hours the client's two-person support team spent Monday and Tuesday matching confused emails to order numbers, because nothing in either system flagged a duplicate on its own. Customers found out before we did. We also paused the next phase of the build — a returns-automation project the client had already signed off on — for three weeks while we rebuilt the sync and, honestly, rebuilt some trust.

A retry doesn't know it's a retry. If nothing checks first, it just does the job again and calls it a success.

What we changed

None of the fixes were exotic.

We ended up writing more about the underlying pattern later — the idempotency check that belongs in front of any webhook-driven workflow, not just this one. This incident is where that lesson came from, even though we didn't say so at the time.

What we tell clients now

Every new integration gets walked through one question before it ships: what happens if this exact step runs twice? If the honest answer is "I'm not sure," it doesn't ship yet. That's a slower way to build. It's also the only way we build automation for clients now, and the slower version is cheaper than the postmortem.

If you've got an integration running today and you can't answer that question for it, that's worth a conversation before it turns into one of these.

— Cole

Got a workflow you're not sure would survive a retry?

30-minute discovery call. We'll walk through what happens when a step in it runs twice.

Book a Discovery Call →