Voice AI

The 207 Problem: Building a Self-Healing Email Batch Sender with Telnyx Edge Agents

If you've ever fired off a large batch email send, you've met the most honest status code in HTTP: 207 Multi-Status. It means exactly what it says — some of this worked, some of it didn't, and now the difference is your problem. Which messages landed? Which failed? How many times has each one been retried? And are you sure the "failed" ones didn't actually go through?

Most systems answer those questions with a spreadsheet, a cron job, and a healthy dose of hope. This post walks through a different approach: the batch campaign itself is a durable, stateful actor that tracks every message, wakes itself up to retry failures with exponential backoff, and leaves behind a complete audit trail. It's built on the Telnyx Email API and the Telnyx Edge Agent SDK — part of Telnyx's AI Communications Infrastructure — and by the end, you'll be able to deploy your own.

The full sample lives here: email-batch-retry-agent. Snippets below are abridged for readability; the repo has the complete implementation.

What the App Does

The demo scenario is a community bank running a fraud-alert campaign: hundreds of cardholders, each notified about a suspicious transaction on their account. Every notice has a deadline — the longer a customer goes without confirming or disputing a charge, the more money is at risk. So "we'll retry later" isn't a plan; it's a liability.

The sample, email-batch-retry-agent, treats the campaign as a first-class running thing:

  • POST /campaigns submits a batch of messages and immediately fires the initial send. The API returns 202 Accepted — the campaign is now running whether you watch it or not.
  • Per-message state is tracked for every message: PENDING, SENT, FAILED, or EXHAUSTED, plus attempt count, last error, and the idempotency key used for each attempt.
  • Self-waking retries: failed messages are retried by tasks the agent schedules for itself, with exponential backoff from 60 seconds up to a 5-minute cap. No cron, no external scheduler, no human babysitter.
  • Durable state: the campaign's full state persists across restarts and crashes. If the platform reboots mid-batch, the campaign picks up exactly where it left off.
  • Audit trail: every campaign is written to a KV store and queryable via GET /campaigns/:id, including per-message attempt counts and idempotency keys.
  • Operator notification: when the campaign reaches a final state, the agent sends an SMS summary to the operator via the Telnyx SMS API.

How It Works

The actor is the campaign

The core design decision: a BatchAgent instance doesn't manage the campaign — it is the campaign. It's born the moment the batch is submitted, it owns the full campaign state, and it survives chaos. That's what the Telnyx Edge Agent SDK provides: a durable actor runtime where persistent storage, queues, and self-waking schedules are primitives, not glue code you write yourself.

Durable state

Every message's fate lives in actor storage:

type MessageStatus = "PENDING" | "SENT" | "FAILED" | "EXHAUSTED";

interface MessageState {
  index: number;
  to: string;
  from: string;
  subject: string;
  text: string;
  status: MessageStatus;
  attempts: number;
  lastError: string | null;
  idempotencyKey: string | null;
}

interface CampaignState {
  campaignId: string;
  messages: MessageState[];
  status: "CREATED" | "RUNNING" | "PARTIAL_FAILURE" | "COMPLETED";
  createdAt: string;
  completedAt: string | null;
}

Because this state is written to ctx.storage, it survives restarts. The agent can crash, the platform can reboot, and the next invocation picks up exactly where the last one left off — including which messages were mid-retry and how many attempts they've burned.

The initial send — and the 207

When a campaign is created, the agent enqueues a sendBatch task. That task calls the Telnyx Email API's batch endpoint and — this is the important part — treats the 207 Multi-Status response as data, not as an error:

async sendBatch() {
  const state = await ctx.storage.get<CampaignState>("campaign");
  const pending = state.messages.filter((m) => m.status === "PENDING");

  // One idempotency key per message, per attempt.
  for (const m of pending) {
    m.idempotencyKey = `${state.campaignId}-${Date.now()}-${randId()}`;
    m.attempts += 1;
  }
  await ctx.storage.put("campaign", state);

  const res = await fetch("https://api.telnyx.com/v2/email_messages/batch", {
    method: "POST",
    headers: {
      Authorization: `Bearer ${ctx.env.TELNYX_API_KEY}`,
      "Content-Type": "application/json",
    },
    body: JSON.stringify({
      messages: pending.map((m) => ({
        to: m.to, from: m.from, subject: m.subject, text: m.text,
        idempotency_key: m.idempotencyKey,
      })),
    }),
  });

  // 207 Multi-Status: the batch was accepted, but each message
  // succeeded or failed independently. Parse it per message.
  const { data } = await res.json();
  data.forEach((r: any, i: number) => {
    const m = pending[i];
    if (r.success) {
      m.status = "SENT";
      m.lastError = null;
    } else {
      m.status = "FAILED";
      m.lastError = r.errors?.[0]?.detail ?? "unknown error";
    }
  });
  await ctx.storage.put("campaign", state);

  if (state.messages.some((m) => m.status === "FAILED")) {
    await ctx.schedule(60, "retryFailed"); // first retry in 60s
  } else {
    await this.finalize();
  }
}

Two details worth calling out:

  1. Idempotency keys are per message, per attempt. Each send attempt gets a fresh key, so the audit trail shows exactly which attempt did what — and a retried message can never be silently double-sent by a duplicate task invocation.
  2. The 207 is parsed, not thrown away. A naive client sees a mixed response and retries the whole batch. This agent updates state for each index individually, so only genuinely failed messages ever go back out.

Self-waking retries

If any messages failed, the agent schedules its own wake-up:

async retryFailed() {
  const state = await ctx.storage.get<CampaignState>("campaign");
  const failed = state.messages.filter((m) => m.status === "FAILED");

  // Exponential backoff: 60s, 120s, 240s … capped at 5 minutes.
  const delaySec = Math.min(60 * 2 ** (failed[0].attempts - 1), 300);

  // Retry only the FAILED indices, each with a fresh idempotency key.
  await this.sendBatch(failed);

  const stillFailing = state.messages.some((m) => m.status === "FAILED");
  if (stillFailing && failed.some((m) => m.attempts < MAX_ATTEMPTS)) {
    await ctx.schedule(delaySec, "retryFailed");
  } else {
    // Anything still FAILED after the final attempt is marked EXHAUSTED.
    await this.finalize();
  }
}

The backoff curve is deliberately boring: 60 seconds, then 2× each round, capped at 5 minutes. Throttling bursts recover in one or two cycles; genuine problems stop burning attempts quickly and get marked EXHAUSTED instead of retrying forever.

And because schedule() is a runtime primitive, the agent sleeps between retries — it isn't polling, and it isn't holding a worker hostage. It wakes, does its job, and goes back to sleep.

The audit trail and the SMS ping

When the campaign reaches a final state, two things happen. First, the full record is written to the KV audit store under campaign:<id>. Second, the agent texts the operator:

async notifyOperator() {
  const state = await ctx.storage.get<CampaignState>("campaign");
  const { sent, failed, exhausted } = summarize(state);

  await ctx.env.TELNYX.sms.send({
    to: ctx.env.OPERATOR_NUMBER,
    from: ctx.env.TELNYX_SENDER,
    text: `Campaign ${state.campaignId}: ${sent} sent, ${failed} failed, ` +
          `${exhausted} exhausted. Audit: GET /campaigns/${state.campaignId}`,
  });
}

GET /campaigns/:id returns the whole story — per-message status, attempt counts, idempotency keys, timestamps:

{
  "campaignId": "campaign-2026-03",
  "total": 100,
  "sent": 98,
  "failed": 0,
  "exhausted": 2,
  "status": "PARTIAL_FAILURE",
  "messages": [ "…per-message records…" ]
}

That response is the whole point. When someone asks "did the fraud alerts go out?", the answer is a URL, not an investigation.

Setup

Clone and install:

git clone https://github.com/team-telnyx/telnyx-code-examples.git
cd telnyx-code-examples/email-batch-retry-agent
npm install

Configure environment variables — copy .env.example to .env and fill in:

VariablePurpose
TELNYX_API_KEYYour Telnyx API key, used for the Email and SMS APIs
TELNYX_SENDERYour verified Telnyx email sender address
OPERATOR_NUMBERThe phone number that receives the SMS campaign summary

Authenticate the Telnyx Edge CLI and run the smoke test:

telnyx-edge auth api-key set "$TELNYX_API_KEY"
npx tsx smoke_test.ts

Deploy:

npm run deploy   # runs `telnyx-edge ship`

Then fire a campaign:

curl -X POST https://your-agent.example/campaigns \
  -H "Content-Type: application/json" \
  -d '{
    "campaignId": "campaign-2026-03",
    "messages": [
      { "to": "+15551234567", "from": "sender@example.com",
        "subject": "Suspicious transaction detected",
        "text": "Review this charge on your account." }
    ]
  }'

You'll get 202 Accepted back immediately. The campaign is now running itself.

Conclusion

Partial failure isn't an edge case in batch messaging — it's the normal case, and the systems that pretend otherwise are the ones that page you at 3 a.m. The durable actor pattern flips the default: state persists, retries schedule themselves, every attempt is idempotent and auditable, and the operator finds out by SMS instead of by spreadsheet.

This sample is a small, complete demonstration of what Telnyx's AI Communications Infrastructure makes practical: the Email API for delivery, the Edge Agent SDK for durability, and SMS for the human loop — one platform, one runtime. Clone the repo, deploy it, and break it on purpose. The campaign will put itself back together.

Ready to build with low-latency voice AI?

Join developers building the future of real-time conversations