The conversational AI stack. Speaks, types, and operates your product.

One core runs text, voice, and visual surfaces. Self-hosts on your hardware, open models included. No per-minute fee, no conversation leaving your infrastructure. This website runs on it.

SELF-HOSTABLESWAPPABLELICENSED SOURCE
Deployment003

One stack, three homes

T1

Cloud API

Managed models over vendor APIs, the runtime in your cloud.

Fastest start, per-minute vendor fees. Every part swaps later without a rebuild.

T2

Private cloud

The full stack inside your VPC, open models on GPU instances you control.

No vendor between you and the weights. Cost is your compute.

T3

Self-hosted

Open weights end to end, on servers in your building.

Internet only where you allow it. Cost is hardware and power, no meter on it.

The system006
01

Runtime

The agent itself. One core, every surface.

  • Orchestration. One core owns state, tools, and conversation flow across every surface.
  • Turn-taking. Interruptions, barge-in, and repair handled in the pipeline.
  • Screen driving. The agent clicks, fills, and navigates your product's UI for the caller.
  • Modality. Voice, text, and screen in one session, switch mid-conversation.
  • Cost gates. Call ceilings, day budgets, and concurrency limits built in.
02

Speech

Recognition and synthesis, open or managed.

  • Recognition. Streaming STT biased to your vocabulary: your product names come out right.
  • Synthesis. Voices cloned or licensed, delivery directed per turn: calm, urgent, warm.
  • Turn detection. Semantic end-of-turn; the reply drafts while the caller still speaks.
  • Languages. The agent detects the caller's language, answers in it, switches mid-call.
  • Noise. Streaming denoise on the input; background chatter does not trip a barge-in.
03

Models

Open or managed models, one interface.

  • Open weights. Open models run on your GPUs; no vendor between you and the weights.
  • Managed. Closed models behind the same interface; the provider swaps by one config line.
  • Tuning. Models fine-tuned on your transcripts: your domain, your terms, your tone.
  • Constrained output. Tool calls grammar-locked to schema; no leaked JSON.
  • Second agent. A background thinker on the idle GPU: same weights, off the critical path.
04

Transport

Every channel into the same pipeline.

  • WebRTC. Calls in the browser and inside your product, LiveKit or raw peer connections.
  • Telephony. SIP trunks and phone numbers terminate in the same agent as the web calls.
  • Relay. Corporate firewalls pass on short-lived TURN credentials, no ports opened.
  • Recovery. Dead connections detected and closed; no orphaned calls holding a slot.
  • One config. Web, in-product, and phone share the same persona and knowledge base.
05

Connectors

Acts in your systems, answers from them.

  • MCP. Client and server; every tool reaches the model through one protocol.
  • Your systems. Bookings, orders, records: the agent operates your product.
  • Knowledge. Retrieval is a connector too: search and citations over your corpus.
  • Trust split. Server tools trusted, client tools sandboxed, kept apart.
  • One entry. Adding a connector is one server entry; the core never changes for it.
06

Reliability

Every call visible and scored.

  • Telemetry. Logs, metrics, traces, and alerts on OpenTelemetry; your tools plug in.
  • Latency split. Response time tiled per stage, mic to voice, each hop on its own meter.
  • Stall naming. A silent turn is logged with a named root cause, per stage, per call.
  • Evals. An LLM judge and transcript replay score conversations from production traffic.
  • Regression. Every change gates on measured conversation quality before it ships.
Reliability005

Every call visible, every stall named

Logs, metrics, traces, dashboards, and alerts ship with the stack, as code. Every stage metered, every stall named, every change scored.

TTFB P5024H
618MS
TTFB P9524H
742MS
Stall rate24H
0.4%
Barge-in24H
3.2%
Response24H · P50 / P95
1.5S0
00:0024:00
Latency split24H · P50
TURN DETECT54MS
STT102MS
LLM118MS
OPENER96MS
TTS244MS
Cost / min24H
$0.06
Tokens / turn24H
148
Day budgetUSED
38%
ConcurrencyLIVE
2/4
GPU utilization24H
100%0
00:0024:00
Tool latency24H · P95
GET48MS
SEARCH96MS
SHOW_VIEW198MS
BOOK142MS
Eval scoreJUDGE · PER RUN
5.00
RUN 01RUN 40
RecognitionNIGHTLY · WER
16%0
RUN 01RUN 40
Stalls by cause7D
RATE_LIMITED4
SERVER_ERROR1
SUPERSEDED3
CANCELLED5
Alert rulesLIVE
TTFB P95 OVER 1.5SOK
STALL RATE OVER 2%OK
LLM 429/5XX RATEFIRING
CALL DROP DETECTEDOK
01 / 0302 / 0303 / 03
Bill of materials006
OPEN · SELF-HOSTED
AI Core
  • RUNTIME
    AGENT ORCHESTRATION · STATE · VIEW DRIVING
  • CONVERSATION
    TURN-TAKING · INTERRUPTIONS
  • SURFACES
    TEXT · VOICE · VISUAL
Speech
  • STT
    NEMOTRON · PARAKEET · WHISPER · MOONSHINE DEEPGRAM · ASSEMBLYAI · GLADIA · SPEECHMATICS
  • TTS
    COSYVOICE · KOKORO · ORPHEUS · PIPER · XTTS CARTESIA · ELEVENLABS · PLAYHT · RIME
LLM
  • OPEN
    GEMMA · QWEN · LLAMA · DEEPSEEK · MISTRAL · PHI · GPT-OSS
  • MANAGED
    GPT · CLAUDE · GEMINI · GROK · NOVA
Transport
  • RTC
    LIVEKIT · WEBRTC RAW DAILY
  • WIRE
    RTVI · WEBSOCKET
  • TELEPHONY
    SIP / PSTN · TWILIO · TELNYX · PLIVO
Reliability
  • TELEMETRY
    OTEL COLLECTOR · LOKI · PROMETHEUS · TEMPO / JAEGER · GRAFANA
  • QUALITY
    EVAL HARNESS · LLM-JUDGE · REGRESSION RUNS · BENCH · TRANSCRIPT REPLAY
Connectors
  • MCP
    CLIENT · SERVER
  • TARGETS
    YOUR SYSTEMS · EXTERNAL APIS