Voice AI

Every Token Your LangChain Agent Emits Can Be Durable, Replayable, and Live — Here's the Shape

Ask a streaming agent to do three things at once — call tools, answer follow-ups, and survive a page refresh — and the tutorials run out. You end up gluing an SSE endpoint, a tool executor, a session store, and a reconnect protocol together yourself.

The LangChain Streaming Agent sample builds that whole stack in one shape, and it works because of what the Telnyx Agent SDK already gives you: a durable message log, a durable event log, and durable tasks — with inference flowing through the pre-authenticated TELNYX binding so no API key ever touches the deployed function.

The Telnyx code example is here:

https://github.com/team-telnyx/telnyx-code-examples/tree/main/langchain-streaming-agent

It is a TypeScript example for Telnyx Edge Compute: a StreamingAgent extends Agent runs a LangChain tool-calling agent, and every token the model emits — plus every tool call and result — streams to the browser the instant it happens.

What This Example Builds

The agent is one durable actor per conversation:

export class StreamingAgent extends Agent<Env, AgentState> {
  @rpc({ description: "Append a user message and run the streaming agent loop" })
  async send(text: string): Promise<{ turn: number }> {
    const state = await this.getState();
    const turn = state.turn + 1;
    await this.messages.add("user", trimmed);   // durable history
    await this.setState({ status: "thinking", turn });
    await this.queue("run");                    // crash-safe agent task
    return { turn };
  }
}

The LangChain wiring is a custom chat model whose tokens come from the TELNYX Inference binding:

export class TelnyxStreamingChatModel extends BaseChatModel {
  async *_streamResponseChunks(messages: BaseMessage[]) {
    const raw = await client.ai.openai.chat.createCompletion({
      model: this.model, messages: mapped,
      stream: true, tools: this.boundTools, tool_choice: "auto",
    });
    for (const chunk of parseStreamedSse(raw)) {
      yield new ChatGenerationChunk({
        message: new AIMessageChunk({ content: chunk.content ?? "" }),
        text: chunk.content ?? "",
      });
    }
  }
}

Three things fall out of that shape for free:

Tokens are durable before they are live. Each streamed delta commits to the agent's event log — this.events.emit("token", { turn, text }) — and the Agent SDK pushes committed events to every attached client. Commit-before-push means a browser that refreshes mid-answer reconnects with resume: true and replays exactly what it missed from its cursor. No missed tokens, no duplicates.

Rapid-fire questions each get their own answer. The run loop drains a backlog: every unanswered user turn (tracked by an answeredThrough message-seq high-water mark) is processed oldest first, each with the history that came before it. Click three prompts in two seconds and you get three answers, in order — instead of the classic bug where only the last question is answered and the rest vanish.

A crash mid-turn reprocesses exactly the unanswered turns. Because answeredThrough only advances after an answer commits, a re-dispatched task picks up precisely where the log says it stopped. The retry logic is the log, not extra code.

The Tool Round-Trip Nobody Warns You About

One detail worth its own section: streaming tool calls break naive model wrappers. The model emits tool calls as deltas — tool_calls[{ index, id, function: { name, arguments: "…partial JSON…" } }] — and on the next round the scratchpad needs them back in full. The sample's toWireMessage handles all three shapes LangChain produces (parsed tool_calls, streaming tool_call_chunks, and the raw additional_kwargs.tool_calls the executor rebuilds scratchpad turns from), so tool_call_id and tool_calls survive every hop. Miss that, and the second model round 400s — usually as a silent retry loop.

Also: AgentExecutor invokes the model per round rather than streaming it, so there are no on_chat_model_stream events to subscribe to. The sample captures tokens at the model layer instead — an onToken hook on the chat model fires per SSE delta, awaited in order, which is what keeps the durable event log clean.

Why It Matters

If your agent runs on the Telnyx Agent SDK, streaming is not an observability project. The history and the progress events already exist, durable and ordered. The sample just points a LangChain agent at them and commits each delta as it happens.

Use cases: support copilots where users fire multiple questions quickly, agents behind flaky mobile connections that must survive refreshes, and any team that wants LangChain's tool ecosystem without owning streaming plumbing.

Run It

git clone https://github.com/team-telnyx/telnyx-code-examples.git
cd telnyx-code-examples/langchain-streaming-agent
npm install
cp .env.example .env
npm run local:dev

Open http://localhost:8787/?session=demo, click a prompt, and watch tokens stream. Deploy with telnyx-edge and the deployed function needs no API key — inference is authenticated by the platform binding. Full setup is in the project README, and the walkthrough (including how to add your own tool) is in GUIDE.md.

Ready to build with low-latency voice AI?

Join developers building the future of real-time conversations