Cadre · Agents for Hire · The delivery layer

Meet the specialists you actually hired.

Not a chatbot box. A room. Named expert teammates — who already know your business — that you see, talk to, and work alongside, live.

🧑‍💼 Named specialists 🗣️ Each with its own voice 🖥️ Two ways to meet them 🧠 They know YOUR business

This blueprint is the how-you-actually-meet-and-work-with-them layer that sits on top of the Cadre venture. Cadre is who the specialists are and why they're a moat. This is the surface where you shake their hand.

Plain English

An AI department you can meet

Imagine hiring a small team of experts — Joe the accountant, Sally in marketing, Tom for legal-intake — except each one has already read your handbook, your past campaigns, your chart of accounts. They don't start from zero every time. They know how you price, how you talk, who your customers are. And you don't just fire questions into a text box — you meet them, hear them answer in their own voice, and watch them work on the same screen you're looking at.

💰

Joe

Bookkeeper. Knows your chart of accounts. Categorises, reconciles, reports — the way your books are kept.

📣

Sally

Marketing. Trained on your brand voice and past campaigns. Drafts that sound like you, not a template.

⚖️

Tom

Legal-intake. Structured intake & triage from your process. Bright line: intake, not legal advice.

👥

Priya

HR / People. Policy-aware onboarding and Q&A, straight from your handbook and SOPs.

The whole point: same engine, N hireable faces — each one steeped in the client's own materials. The difference between "an impressive demo" and "a colleague" is context and role, and that's exactly what you're meeting here.

Two ways to work with them

Two delivery tiers

There are two surfaces where you meet your hired specialists. One is light and always-on. The other steps up to a full shared live room for a proper working session. Same specialists, same brains, same voices — different venues.

Tier 1 · always available

The Workbench

A web page of your own — avatar tiles down one side, the shared output in the middle, a chat-and-voice bar at the bottom. Open it anytime, talk to any specialist, watch the work land on the stage. The lighter, everyday surface.

  • Avatar tiles you click to bring a specialist forward
  • Type or speak; each replies in its own voice
  • Shared output stage in the middle
  • Always on — no meeting to schedule
Tier 2 · stepping up

The Integrated Zoom Room

When it's time for a real working session with your team and clients across locations, the specialist joins a Zoom call, shares the working screen, and speaks into the room. Humans take turns at the wheel; the agent drives underneath the whole time.

  • Everyone in one shared room, across miles & devices
  • Humans take turns via Zoom remote-control
  • Agent's voice piped into the call so all hear it
  • The wow version — a live teammate in the meeting

Think of the Workbench as your specialists' office you can drop into any time, and the Zoom room as booking them into a live meeting when the work needs everyone in the same place at the same time.

Tier 1 · the in-between

The Workbench, up close

Our own web page — workbench-style. Avatar tiles are positioned on screen, the shared output sits in the middle, and you interact by text and voice. Each specialist has its own persona and voice. Playwright drives the interactivity underneath, so tiles, the stage, and the work all move as one live surface.

Your Workbench — Acme Co. · 4 specialists on the bench

Your bench

💰JoeBookkeeper🔊
📣SallyMarketing
⚖️TomLegal-intake
👥PriyaHR / People

Shared stage

Joe is drafting · Q3 P&L

Talk to Joe

YouCan you pull together the Q3 P&L the way we did last quarter?
🔊 JoeOn it — same format as Q3 last year, using your chart of accounts. Draft's on the stage.
Type to Joe…🎙️

A representative sketch of the Workbench surface — avatar tiles, a shared output stage, and a text-plus-voice bar. Icon avatars first; video avatars are a deferred later layer.

🖼️

Avatar tiles

Each specialist is a persona tile. Click to bring them forward; the active one lights up and the stage follows their work.

🎬

Shared stage

The middle is the shared output — the document, the report, the draft — updated live as the specialist works.

🗣️

Text & voice

Type or speak. Every reply is spoken back in that specialist's own voice, so the team feels like people, not one bot.

How the Workbench is wired (the specifics)
  • The page: our own single-page surface, served from Cloudflare. Avatar tiles, a shared-output stage, a text+voice composer. No app to install.
  • The hands: Playwright drives a real browser session per specialist underneath — clicking, filling, navigating, producing the artifact that appears on the stage. The interactivity is real automation, not a canned animation.
  • The voice: each reply is spoken with a per-persona voice (Piper / streaming TTS). Joe sounds like Joe; Sally sounds like Sally.
  • The brain: a model-agnostic agent CLI runs the live turn. Which model is a swappable, margin-tunable choice — the persona doesn't change when the brain does.
  • The memory: Zep carries each specialist's persistent memory of your business across sessions, so the context compounds instead of resetting.
  • Metering-clean: every turn is a live, human-in-the-loop dispatch — never a headless per-user bot running on its own.
Icon avatars now · video avatars later (deferred, flagged as extra cost)

Start with icon avatars — a coloured tile with a face/emblem per specialist. Cheap, instant, works everywhere, and already conveys "a team of people."

Video avatars are a later layer — animated talking-head avatars that lip-sync the voice. They are a genuine step up in presence but come with real added cost and complexity (per-avatar rendering, licensing, latency). Deferred on purpose: this is an extra expense layer to switch on once the core surface earns it, not a launch requirement.

Tier 2 · the wow version

The Integrated Zoom Room

This is the full live-shared-session room — the "Zoom with a computer we own on the call" architecture. Your specialist joins a Zoom, shares the working screen, and speaks into the room. Your team and clients — across locations — take turns at the wheel via Zoom remote-control, while the agent is always driving the browser underneath via Playwright. Its voice is routed into the call so everyone hears it.

LIVEAcme Co. · Working session — Q3 close🔊 Joe is speaking
💰Joespecialist · drivingunderneath
OOmarAcme, HQhas the wheel
JJackAcme, remote🎤
CClientwatching live🔇

One person holds the Zoom wheel at a time; the agent holds control underneath continuously. Everyone watches the same live screen and hears the specialist speak.

🖥️

Playwright = the hands

The agent drives a real browser on the VM — the constant hand on the wheel, whatever happens on top.

📹

Zoom = the room

The VM joins the call and screen-shares that browser. The agent's screen becomes the meeting's shared screen.

🎚️

Humans take turns

Zoom's Request / Give Control hands the wheel to one person at a time; they act, then hand it back.

🔊

Voice into the mic

The agent's speech is routed via VB-Cable into Zoom's microphone input — so the whole room hears it live.

The key idea: whoever holds the Zoom control on top, the agent still holds control underneath through Playwright. Humans take turns; the agent is the constant hand beneath them. This is the deep architecture — we reuse it, we don't rebuild it.

📹 Open the full live-session blueprint →
The Zoom-room architecture in five moves (deep detail)
MoveWhat happens
1 · HandsPlaywright drives a real browser on the VM — the automation layer. The agent's hands stay on the wheel continuously, whatever happens on top.
2 · The roomThe VM joins a Zoom call and screen-shares that browser window. The agent's screen is now the meeting's shared screen.
3 · Taking turnsZoom's Request / Give Remote Control lets one human take the wheel for a turn, then hand it back. The agent keeps working underneath.
4 · The voiceText-to-speech plays out of the VM and is routed into Zoom's mic via VB-Cable (a virtual audio cable). Everyone on the call hears the agent speak live. Bridge / fast-chat text runs in parallel.
5 · More teammatesA second VM joins the same Zoom the identical way — its own browser, its own voice cable. Now two specialists work the room alongside the humans.

Full deep-dive with the layered diagram lives at the live-session blueprint — this page links to it rather than re-explaining it from scratch.

Where a specialist comes from

Creating a specialist

Before you can meet Joe, Joe has to be made for you. That starts with an onboarding interview, runs your knowledge base through ingestion, and ends with a provisioned specialist ready on your bench.

1Onboarding interview

An agent interviews you — role, tone, what "good" looks like. This is the made-for-you moment and the literal config.

2Ingest your KB

Your SOPs, brand voice, pricing, chart of accounts, past work — read in and indexed (RAG). The context that makes it yours.

3Provision the specialist

Persona + KB + voice bound into a named runtime employee. Joe now exists on your bench.

4Meet & work

Joe appears on the Workbench and can join a Zoom room. He already knows your business on turn one.

The four orthogonal layers — the core IP

Persona ⊥ Agent ⊥ Model ⊥ Voice. Any employee wears any hat, on any brain, in any voice. Independent by design — that's the product feature and the reason the client is locked only to their own context.

🎭Persona
the job + your KB
🤖Agent
the named runtime employee
🧠Model
the swappable brain
🗣️Voice
the face
What each layer means (and why they're independent)
LayerWhat it isWhy orthogonal
PersonaThe role plus the client's KB — "Bookkeeper who knows Acme's chart of accounts."The durable value. Carries across any brain or voice.
AgentThe named runtime employee — "Joe" — the thing you actually put to work.A stable identity even as the model underneath changes.
ModelThe swappable brain — Claude, GPT, Gemini, Grok. A margin-tunable input.Commoditises monthly; never the lock-in. Swap freely.
VoiceThe face — the per-persona voice each specialist speaks in.Recognisable team, chosen not cloned (voice-cloning stays gated).

Why this matters commercially: because these are independent, the client is locked only to their own accumulated context — not to a model, a vendor, or a voice. The interchangeability is the feature; the context is the moat.

Why this wins

The strategy

🏰

Moat = your own context

The durable value isn't the model — models commoditise monthly. It's the accumulated, proprietary context of your business, encoded into a role. The longer a specialist works with you, the more it knows, the higher the switching cost. The moat compounds.

🔄

Model-agnostic

The brain is a swappable margin input, not a lock-in. Run any specialist on whichever model fits — quality, cost, or speed. When a cheaper or better model lands, switch it; the persona and its context don't change.

🎯

vs. generic chatbots

A generic "marketing agent" is a commodity a $20 lab replicates. One steeped in your company's knowledge is not. The gap isn't model capability — it's context and role. That's the whole wedge.

The load-bearing claim: an agent that has read your handbook, your past campaigns, your chart of accounts is an employee, not a tool. The delivery experience on this page — meeting them, hearing them, watching them work — is what makes that difference felt, not just claimed.

Under the hood

What we'd use

The delivery experience is a stack of proven, mostly metering-clean plumbing. Here's what does what — with dropdowns for the depth.

🖥️Playwright
The hands. Drives a real browser per specialist — the constant automation layer, on the Workbench and under the Zoom share.
🗣️Piper / streaming TTS + VB-Cable
The voice. Per-persona speech; on Zoom, routed into the call's mic via VB-Cable so the room hears it.
📹Zoom
The room. Screen-share of the driven browser, turn-based remote-control hand-off, and the audio mix everyone hears.
📚RAG / KB ingestion
The context. Reads and indexes the client's SOPs, brand voice, pricing, past work — what makes a specialist theirs.
🧠Model-agnostic agent CLIs
The brains. Named personas run on any lab's CLI — the swappable, margin-tunable input.
💾Zep
The memory. Persistent per-specialist memory of the business, so context compounds across sessions.
☁️Cloudflare
Delivery. Serves the Workbench, tunnels the sessions, and scopes access per client.
Playwright — the hands (detail)
  • Drives a real browser session on the VM — clicking, filling, navigating, producing the artifact on the stage or the shared screen.
  • On the Workbench: powers the live interactivity of the tiles and stage.
  • In the Zoom room: the constant hand underneath, regardless of who holds the Zoom control on top.
  • Pure plumbing — no model call. The brain is always the live interactive session.
Voice — Piper / streaming TTS + VB-Cable → Zoom (detail)
  • Each persona has its own voice; replies are spoken, not just printed.
  • On the Workbench, audio plays in the page. In the Zoom room, VB-Cable routes that speech into Zoom's microphone input so the whole room hears it.
  • Honest limit: TTS has a short turnaround — near-real-time for quick exchanges, longer for deep answers. Live enough to feel like a teammate; not instant like a phone call.
  • Voice selection is fine; voice cloning stays behind the hard consent/labeling gate.
Zoom — room, turn-based control, audio mix (detail)
  • The VM joins the call and screen-shares the driven browser — the shared space everyone sees.
  • Request / Give Remote Control hands the wheel to one human at a time — cooperative hand-off, not many cursors at once.
  • The agent's voice is mixed into the call audio via the virtual cable.
  • Deep architecture: the live-session blueprint.
RAG / KB ingestion, model CLIs, Zep, Cloudflare, relay (detail)
  • RAG / KB ingestion: the onboarding step that reads the client's docs and makes them retrievable to the specialist. "Train it on our docs" is now routine, not a research project.
  • Model-agnostic agent CLIs: a persona can run on any major lab's CLI — the brain is a swappable, margin-tunable input.
  • Zep: persistent semantic memory per specialist, so the accumulated context compounds (and is the moat).
  • Cloudflare: serves the Workbench page, tunnels the live sessions, and scopes access per client (link-gated or per-person via CF Access).
  • Persona / relay substrate: the hardened relay that drives real interactive agent CLIs by name — the light, proven path replacing the earlier OOM-crashing multi-agent build.
No hand-waving

Honest limits

What this delivery experience genuinely can and can't do today — stated plainly, because the honesty is the credibility.

🔁

Turn-based, not simultaneous

One controller holds the wheel at a time — cooperative hand-off, not many live cursors at once. True simultaneous multiplayer is a different, much harder architecture — and it isn't needed for turn-taking work.

⏱️

Voice has turnaround

It's text-to-speech with a short delay — near-real-time for quick exchanges, longer for deep answers. Live enough to feel like a teammate; not instant like a phone call.

🎥

Video avatars = added cost

Talking-head avatars are a real step up in presence but carry genuine cost and latency. Deferred on purpose — icon avatars first, video switched on only once the core earns it.

📊

The real gate: metering / throughput

Every specialist must be a live-session-dispatched instance with a human in the loopnever a headless per-user bot. So the true scale question isn't "how clever is the agent" — it's how many simultaneous live sessions one router / box sustains. That's what decides how many clients you can serve at once.

Why the metering line is the commercial gate, not a footnote: the whole model stays clean only if there's always a human-driven session in the loop. That makes serial live-session throughput — not model capability — the ceiling on scale. Design around it, and the economics hold; pretend it away with headless fan-out, and both the ethics and the metering break.

The portfolio

How it connects

This delivery layer doesn't stand alone — it's the live surface for a venture and a stack that already exist.

◆ Cadre

The venture underneath — who the specialists are, why the client-KB moat is the wedge, the tiers and go/no-go. This page is Cadre's delivery experience.

Cadre hub →

🛍️ AgentAIShop

The build-and-sell + marketplace side. The same onboarding-interview pattern that provisions a Cadre specialist is where AgentAIShop's creator flow came from. Build here, list there.

AgentAIShop → · Marketplace →

🔨 The Forge

The venture workshop this blueprint was forged in — where proto becomes a real, hosted, editable artifact. Cadre and this delivery layer are Forge outputs.

The Forge →

💾 Sidecar memory-moat

The same context-lock-in logic Sidecar documents for individuals, lifted to the company. Zep is the shared spine; the compounding context is the moat on both.

The one line

Hire the specialists who already know your business — then meet them.

An AI department you can see, talk to in their own voices, and work alongside — on a light everyday Workbench, or stepping up into a full live Zoom room. Same specialists, same context, same moat.

◆ The Cadre venture 📹 The live-session architecture
◆ Cadre hub