Drain the webhook retry backlog

Node · Node · intermediate · greenfield

Adds the retry-backlog drain job: it re-signs and re-sends every pending webhook delivery concurrently and reports a per-delivery outcome, with the existing per-delivery timeout and error isolation carried over. Ran it against a staging backlog of a few dozen deliveries and everything went out in one quick burst with correct results.

The webhook subsystem of a payments platform. The drain job runs after provider outages, when the backlog is at its largest; each delivery is an HTTPS POST to a customer endpoint. Node 20, global fetch.

Requirements

Files touched

--- src/webhooks/drain.js
+const crypto = require('node:crypto');
+
+/**
+ * Webhook retry backlog drain. After an outage the backlog can hold
+ * thousands of pending deliveries; this job re-sends them all.
+ */
+
+function sign(body, secret) {
+  return crypto.createHmac('sha256', secret).update(body).digest('hex');
+}
+
+async function deliverOne(delivery) {
+  try {
+    const res = await fetch(delivery.url, {
+      method: 'POST',
+      headers: {
+        'Content-Type': 'application/json',
+        'X-Webhook-Signature': sign(delivery.body, delivery.secret),
+      },
+      body: delivery.body,
+      signal: AbortSignal.timeout(5000),
+    });
+    await res.body?.cancel(); // discard the body so the connection is released
+    return { id: delivery.id, ok: res.ok, status: res.status };
+  } catch {
+    return { id: delivery.id, ok: false, status: 0 };
+  }
+}
+

Review this PR

Node practice