Voice AI

Building Durable Multi-Agent Workflows with Telnyx Edge Actors

When Your 3 AM Batch Job Crashes, You Need More Than Retries

Imagine you're running a clinic-network transcription service. Every night at 3 AM, your system kicks off dozens of speech-to-text jobs on patient recordings. The power goes out. When you come back online, you have no idea which files were transcribed, which failed, and which never started. Do you restart everything? Do you manually check each file? Or do you build a system that remembers exactly where it left off?

That's the problem we set out to solve with the Sub-Agent Orchestrator Actor — a durable parent actor that manages parallel transcription jobs, tracks their lifecycle, and resumes interrupted runs without redoing finished work. Built entirely on @telnyx/edge-runtime 0.15.2.

What the App Does

The orchestrator is a persistent parent actor that:

  1. Spawns one child actor per audio file — each child is a persistent actor with its own state, not a transient task
  2. Tracks each child's lifecycle — PENDING → RUNNING → COMPLETED → FAILED
  3. Persists results to KV before reporting — every child writes its outcome to KV before telling the parent it's done
  4. Resumes interrupted runs — if the parent or any child crashes, re-posting the same jobId picks up exactly where it left off
  5. Self-cleans on completion — when all children are done, the parent destroys itself

The result is a system that's resilient to crashes at any level. A power event at 3 AM doesn't mean re-transcribing everything — it means re-spawning only the workers that never reported.

How It Works

The Parent Actor: OrchestratorAgent

The parent actor IS the workflow. It owns the job state, spawns children, tracks their lifecycle, accumulates results, and self-cleans when the job completes.

// Spawn one child per audio file
for (const file of files) {
  const childName = `${jobId}:${file.id}`;
  await this.spawn(TRANSCRIBER_ACTOR, childName, { file });
}

// Track children lifecycle
const children = await this.children();
for (const child of children) {
  const status = await this.getKV(`job:${jobId}:file:${child.fileId}`);
  // Update scorecard based on KV state
}

The Child Actor: TranscriberAgent

Each child is a persistent actor that:

  1. Receives its payload via assign()
  2. Calls the Telnyx AI transcription endpoint
  3. Writes its outcome to KV first
  4. Reports completion back to the parent via the PARENT binding
// KV-first write order — this is critical
await this.putKV(`job:${jobId}:file:${fileId}`, {
  status: 'COMPLETED',
  text: transcription.text,
  attempts: this.attemptCount
});

// Only after KV is written, report to parent
await this.reportComplete({ fileId, text: transcription.text });

KV as the Source of Truth

KV is the authority for job state. Every child persists its per-file outcome to KV before reporting to the parent. The parent derives progress from KV, so a crash anywhere — child or parent — is recoverable by re-posting the same jobId.

The KV schema:

  • job:{jobId} — full job record + scorecard
  • job:{jobId}:file:{id} — per-file outcome (child-written)
  • job:{jobId}:attempt:{id} — attempt ledger (parent-written)

Recovery Without Redoing Work

When a job is re-posted after a crash, the parent's reconcile() method:

  1. Reads all existing KV entries for the job
  2. Identifies which files are already COMPLETED or FAILED
  3. Re-spawns only the children that never reported
  4. Bounds re-spawns by MAX_CHILD_ATTEMPTS

This means if 47 out of 50 files were transcribed before a crash, only 3 children are re-spawned.

Stuck-Child Detection

A watchdog timer detects children that are RUNNING but haven't reported within STUCK_TIMEOUT_SECONDS. These are treated as never-reported and re-spawned (if under the attempt limit).

Setup

Prerequisites

  • Node.js 18+
  • Telnyx CLI
  • A Telnyx account with API key access

Environment Variables

VariableRequiredDescription
DEMO_MODEYestrue (default) mocks transcription; false enables live calls
MOCK_AUDIO_URLSYesComma-separated audio file list
TELNYX_API_KEYLive modeTelnyx API key for REST calls
OPERATOR_NUMBERLive modePhone number for SMS notifications
TELNYX_SENDERLive modeTelnyx number that sends SMS
TRANSCRIPTION_MODELNoSpeech-to-text model (default: distil-whisper/distil-large-v2)
STUCK_TIMEOUT_SECONDSNoWatchdog window (default: 300)
MAX_CHILD_ATTEMPTSNoMax re-spawn attempts per file (default: 3)

Running Locally

# Clone the repository
git clone https://github.com/team-telnyx/telnyx-code-examples.git
cd telnyx-code-examples/sub-agent-orchestrator-actor

# Install dependencies
npm install

# Set environment variables
export DEMO_MODE=true
export MOCK_AUDIO_URLS="https://example.com/audio1.mp3,https://example.com/audio2.mp3"

# Start the local runtime
npx telnyx dev

Then visit http://localhost:8080/ to see the operator console.

Deploying to Telnyx Edge

# Deploy
npx telnyx deploy --name sub-agent-orchestrator

# Set production secrets
npx telnyx secrets set TELNYX_API_KEY
npx telnyx secrets set OPERATOR_NUMBER
npx telnyx secrets set TELNYX_SENDER

Why Telnyx

Telnyx provides the AI Communications Infrastructure that makes this sample possible — a single platform where durable edge actors, AI inference, and SMS notifications work together. The Telnyx Edge Runtime gives you persistent, stateful actors that can spawn and manage child actors, while the Telnyx API handles operator notifications and AI-powered transcription summaries. No external orchestration services, no glue code — just one platform for the entire workflow.

Conclusion

The Sub-Agent Orchestrator Actor demonstrates how Telnyx Edge Runtime's actor model can be used to build genuinely durable, distributed workflows. By treating each transcription job as a persistent child actor and using KV as the source of truth, we've built a system that survives crashes without wasting compute on redundant work.

The key insights:

  • KV-first writes ensure durability before reporting
  • Persistent child actors (not transient tasks) enable lifecycle tracking
  • Reconciliation on resume means only unfinished work is re-spawned
  • Self-cleaning prevents resource leaks

This pattern isn't limited to transcription — it applies to any batch processing workflow where reliability and efficiency matter. Whether you're processing medical records, financial documents, or media files, the orchestrator pattern gives you crash-resilient parallelism with honest progress tracking.

Try it yourself in the GitHub repository.

Ready to build with low-latency voice AI?

Join developers building the future of real-time conversations