Most incident systems store an outage in one database, customer impact in another, notifications in a queue, and the RCA in a document store. Every service understands one slice of the incident, but nothing owns the whole lifecycle.
The network-incident-agent sample takes a different approach: the actor is the outage. One durable NetworkIncidentAgent instance owns one incident ID from detection through investigation, restoration, resolution, RCA generation, and a later recurrence check.
The complete sample is available at https://github.com/team-telnyx/telnyx-code-examples/tree/main/network-incident-agent.
What is a network incident agent?
A network incident agent is a durable software actor whose identity is an incident ID. Sending work for INC-EDGE-042 to the actor namespace always reaches the same logical entity. That entity owns status, severity, affected services, customer impact, notification counts, resolution metadata, and recurrence state.
This is different from making an AI chat session about an outage. The conversation is temporary; the incident is the durable object.
Alert or operator action
|
v
NetworkIncidentAgent("INC-EDGE-042")
|-- durable incident state
|-- affected customers in KV
|-- lifecycle events in SQL
|-- severity via AI Inference
|-- customer SMS via Messaging
|-- incident context via Voice
|-- RCA JSON in CloudFS
`-- recurrence check via schedule()
A durable incident lifecycle
The actor accepts only valid transitions:
detected -> investigating -> restoring -> resolved -> closed
^ |
`---------------'
Each transition updates durable state and writes an event to the actor's embedded SQL database. The timeline is therefore operational data rather than a reconstructed application log.
When the incident is resolved, the actor writes a structured RCA document to CloudFS using a temporary file and atomic rename. It then schedules a delayed recurrence check. If the incident remains resolved, the check closes it. If its state regresses, the actor records a recurrence.
Keep customer data out of the UI
Affected phone numbers are stored in KV. Browser snapshots expose masked values such as •••0100, and timeline messages avoid raw phone numbers and message bodies. This matters for an operations dashboard that may be screen-shared or recorded.
The sample also separates two execution modes:
- Demo mode exercises durable state, KV, SQL, CloudFS, lifecycle transitions, and scheduling while simulating SMS delivery.
- Live mode uses Telnyx Inference and Messaging and requires real account configuration plus opted-in destinations.
The application never reports simulated delivery as live delivery, and failed live requests surface as visible errors.
Proactive SMS and incident-aware calls
The actor sends status messages through the zero-credential Telnyx binding:
await this.env.TELNYX.messages.send({
from,
to,
text: params.message,
});
For voice, an inbound call.initiated webhook carries the incident ID in client_state. The application resolves the corresponding actor, asks it for current incident context, answers the call, and speaks the update through Call Control.
That means a customer receives the same current state whether they read an SMS or call for an update.
Run the safe demo
git clone https://github.com/team-telnyx/telnyx-code-examples.git
cd telnyx-code-examples/network-incident-agent
cp .env.example .env
npm install
npm run build
mkdir -p /tmp/network-incident-cloudfs
CLOUDFS_MOUNT_PATH=/tmp/network-incident-cloudfs npm start
Open the URL printed by telnyx-edge, normally http://localhost:8787. Leave Send real Telnyx SMS unchecked and run the guided incident.
The dashboard shows severity assessment, affected services, notification counts, each lifecycle transition, the SQL timeline, masked customers, the voice response preview, the generated RCA, and the recurrence task.
You can also validate the API flow:
DEMO_BASE_URL=http://localhost:8787 npm run smoke
Where this pattern fits
The same durable-entity pattern works for more than network operations:
- Carrier and regional service outages
- Enterprise incident communications
- IoT fleet degradation
- Security incident coordination
- Planned maintenance windows
- Customer-specific service-impact cases
The important design decision is consistent: make the long-lived business entity the actor. Let communication channels and AI reasoning operate inside that durable lifecycle instead of becoming the lifecycle themselves.
Production considerations
Before production use, protect operator endpoints with authentication and authorization, validate webhook signatures at ingress, use idempotent command IDs for at-least-once delivery, configure a real KV namespace and CloudFS mount, and restrict live delivery to approved destinations.
The sample is deliberately small, but the architecture is real: one incident, one durable actor, one source of current operational truth.