# Connecting Your Own AI Bot to a Phone Call — Streaming (BYO-AI) Guide

**Audience:** developers integrating their own voice AI (Gemini, GPT Realtime, a
custom model — anything) into a live phone call, through the **Streaming** channel
(מודול "סטרימינג").

**Model:** you expose a **WebSocket endpoint**. For every call, our PBX connects
**out** to it, streams the caller audio to you, plays your AI's audio back to the
caller, and lets your AI drive **real call actions** (transfer, hangup) over a JSON
control channel. You run and pay for the AI; we provide the telephony channel,
billed per channel-minute.

> Built-in agents (מודול "בוט AI") run on our Gemini and need none of this. This
> guide is **only** for bringing your *own* AI via a Streaming channel.

```
caller ──phone──> PBX ──8k PCM (binary WS)──> your WebSocket ──> your AI
caller <─phone── PBX <─8k PCM (binary WS)──── your WebSocket <── your AI
                  PBX <── JSON commands ────── your AI   (transfer / hangup / …)
                  PBX ──> JSON events ───────> your AI   (start / dtmf / hangup)
```

---

## 1. Create the channel (one-time, in the management UI)

**מודולים → סטרימינג → ＋ ערוץ סטרימינג:**

| Field | Meaning |
|---|---|
| נקודת קצה (WebSocket) | your `wss://…` endpoint — where we connect for each call |
| סוד אימות (Bearer) | optional shared secret; if set we send it as a Bearer token |

Save → the channel gets a **token** (its stable id). Then attach it to a שלוחה in
the IVR tree via a **סטרימינג** node that references this token. When a caller
reaches that node, we open a WebSocket to your endpoint for that call.

---

## 2. Connection

For every call we open one WebSocket to your configured URL:

```
GET wss://your-ai.example.com/voice
Authorization: Bearer <your-secret>      (only if you set a secret)
```

Accept it and keep it open for the duration of the call. One socket = one call.

---

## 3. Audio format

- **Raw PCM, 16-bit signed, little-endian, mono, 8000 Hz** (telephony).
- Carried as **binary** WebSocket frames, both directions.
- From us: one frame ≈ **320 bytes every 20 ms** (8000 Hz × 2 bytes × 0.02 s).
- To us: **any chunk size**, but **send at ~real-time**. We buffer ~10 s of your
  audio; sustained faster-than-real-time output will back-pressure (and waste your
  AI's tokens). Resample to/from your model's native rate on your side
  (e.g. 8k↔16k in, 24k↔8k out).

---

## 4. Events we send you  (server → your bot)

A **binary** frame is caller audio. A **text** frame is a JSON event:

```jsonc
// sent once, immediately after connect — your cue to init the AI session
{"type":"start","callId":"<uuid>","caller":"05XXXXXXX","system":"07XXXXXXX",
 "token":"<channel-token>","format":"pcm16;rate=8000;ch=1"}

{"type":"dtmf","digit":"5"}    // the caller pressed a keypad key
{"type":"hangup"}              // the caller hung up — close the socket and stop
```

| Field on `start` | Meaning |
|---|---|
| `callId` | unique id for this call (use it to correlate your own logs) |
| `caller` | the caller's phone number (caller ID), if available |
| `system` | the PBX system number the call belongs to |
| `token`  | the Streaming channel's token |
| `format` | always `pcm16;rate=8000;ch=1` for now |

---

## 5. Actions your bot sends us  (your bot → server)

A **binary** frame is your AI's voice (played to the caller). A **text** frame is a
JSON **command**; the `type` field selects the action.

| `type` | fields | terminal? | effect |
|---|---|:--:|---|
| `transfer_extension` | `target` = IVR path | **yes** | route the call to an IVR extension/path |
| `transfer_phone` | `target` = phone number | **yes** | dial an external phone number |
| `transfer_queue` | `target` = queue id | **yes** | route the call to an agent queue |
| `hangup` | — | **yes** | end the call |
| `clear` | — | no | barge-in: drop the audio we still have queued to play |
| `save_answer` | `field`, `value` | no | log a survey answer (shows in the channel logs + Excel export) |

**Terminal vs non-terminal:** a *terminal* command ends the streaming session — we
hand the live call back to the PBX, which executes the action, and the WebSocket
closes. `clear` and `save_answer` are *non-terminal*: the conversation continues.
Send **at most one** terminal command per call (the first one wins; anything after
is moot because the session is already ending).

**Closing words first:** after a terminal command we **drain your queued audio for
up to ~6 s** before executing — so the caller hears your AI's final sentence
("I'm transferring you now…") before the transfer/hangup happens. Send the command
right after you start that sentence; you don't need to wait for it to play out.

**Where the action runs:** terminal actions execute on the **live channel** through
the exact same PBX/AGI engine the built-in bot uses — so an extension transfer
navigates the real IVR tree, a phone transfer dials out with the system's caller
ID, etc.

### The actions in detail

```jsonc
// → route to an IVR extension/path. target is an IVR path, same format as the
//   extension tree (e.g. "/2" = option 2 of the main menu, "/sales/agent").
{"type":"transfer_extension","target":"/2"}

// → dial an external phone number and connect the caller to it.
{"type":"transfer_phone","target":"03XXXXXXX"}

// → route the caller into an agent queue (by queue id).
{"type":"transfer_queue","target":"<queue-id>"}

// → end the call.
{"type":"hangup"}

// → barge-in. Call this the moment the caller starts talking over your AI, to
//   immediately stop whatever you've queued for playback. (Non-terminal.)
{"type":"clear"}

// → record a structured answer for this call. Appears per-call in the channel
//   logs and in the survey Excel export, keyed by `field`. (Non-terminal.)
{"type":"save_answer","field":"id_number","value":"12345678"}
```

> ⚠️ `target` must come from the call flow / your AI's own logic — **never** put a
> number the caller dictated (their ID, credit card, etc.) into a `transfer_*`
> target. That would dial it. Capture such values with `save_answer`.

---

## 6. Call lifecycle

```
connect            → (verify Bearer) → init your AI session
recv {type:start}  → you now have callId / caller / token
recv binary        → caller audio → feed your AI
send binary        → your AI's voice → caller hears it
recv {type:dtmf}   → caller pressed a key → react if you want
send {type:clear}  → caller barged in → drop your queued audio
send {type:save_answer,…} → log a field (repeatable)
…
send {type:transfer_extension,…}  → say your closing line, then send this
  ← we drain ≤6 s of your audio, execute on the live channel, close the socket
        — or —
recv {type:hangup} → caller hung up → close the socket and stop
```

---

## 7. Minimal example (Node.js + `ws`)

```js
import { WebSocketServer } from "ws";

const SECRET = process.env.STREAM_SECRET; // same value you set in the UI
const wss = new WebSocketServer({ port: 8080, path: "/voice" });

wss.on("connection", (ws, req) => {
  if (SECRET && req.headers["authorization"] !== `Bearer ${SECRET}`) {
    ws.close(1008, "unauthorized");
    return;
  }
  let call = null;

  ws.on("message", (data, isBinary) => {
    if (isBinary) {
      // caller audio: 8kHz PCM16 mono LE. Resample to your model, feed it in.
      onCallerAudio(data);
      return;
    }
    const msg = JSON.parse(data.toString());
    switch (msg.type) {
      case "start":  call = msg; initAI(call); break;       // callId, caller, token…
      case "dtmf":   onDigit(msg.digit); break;
      case "hangup": ws.close(); break;
    }
  });

  // your AI produced audio → downsample to 8kHz PCM16 and send as binary
  function speak(pcm8k) { ws.send(pcm8k, { binary: true }); }

  // your AI decided to act → send a JSON command
  function transferToAgent() { ws.send(JSON.stringify({ type: "transfer_extension", target: "/2" })); }
  function logAnswer(f, v)   { ws.send(JSON.stringify({ type: "save_answer", field: f, value: v })); }
  function bargeIn()         { ws.send(JSON.stringify({ type: "clear" })); }
});
```

(Single-writer note: if your stack can write to the socket from multiple async
tasks, serialize the writes — one writer per connection.)

---

## 8. Billing

Streaming (BYO-AI) is billed **per channel-minute**. You pay your AI provider
directly for tokens. (Built-in agents are billed by AI tokens instead.)

---

## 9. Testing & troubleshooting

- **No connection / immediate close** — check the WSS URL and that the Bearer
  secret matches the channel config exactly.
- **Caller hears nothing** — confirm you're sending **binary** frames of 8kHz
  PCM16 mono LE; text frames are treated as control, not audio.
- **Robotic / chipmunk audio** — sample-rate mismatch; we are strictly 8000 Hz.
- **Transfer "does nothing"** — make sure it's a **terminal** command with a valid
  `target` (an existing IVR path / reachable number / real queue id), and that you
  finished (or started) your closing sentence so the ≤6 s drain has something to
  play.
- **Survey columns empty** — you must send `save_answer` per field; nothing is
  inferred from the transcript.
