No black box.
Here's how it thinks.
Every employee reasons step by step, remembers context across tasks, and passes each action through a gate that refuses what's off-domain, pauses what's risky, and runs the rest — at the autonomy level you set.
It reasons, then checks its own work.
A chatbot guesses the next word and ships the first draft. Our employees draw on 200+ reasoning and skill methodologies — chain-of-thought, tree-of-thought, self-critique loops — so they work a problem in stages and catch their own errors before the work reaches you.
Multi-stage reasoning
Each step is validated before the next one runs.
Self-critique
A built-in critic re-reads the work and flags weak spots.
Shows its reasoning
You see the why — methodology, confidence, and risk — not just the answer.
44 specialists
A domain expert per role, not one generalist stretched thin.
Memory that persists across tasks.
An 8-tier memory architecture backed by Postgres and an optional knowledge graph, so an employee remembers what happened yesterday — and what another employee already learned. Context is shared, not re-explained.
What happened, and when — the running record of work done.
What it knows — your facts, docs, and domain knowledge.
How it does the job — the steps and skills it has learned.
One gate, no exceptions. You set the limits.
Before an employee acts, the request passes a single decision point: is it in this employee's lane, free of injection, and safe to run? Off-domain or manipulated requests are refused. Risky ones pause for your approval. Everything else runs.
Refuse
Off-domain or out-of-scope requests are turned down before anything happens.
Pause
High-risk actions wait in a human-in-the-loop queue with full context.
Run
Safe, in-scope work proceeds — in co-pilot it drafts for you; in auto-pilot it executes.
Co-pilot (they draft, you approve) → auto-pilot (they execute) is a switch you flip per employee. A kill switch stops any of them instantly.
We measure what we can — and don't invent what we can't.
There's no honest way to print one “accuracy score” for an AI employee — so we don't. Here is what we actually check.
Refusal, before generation
A scope gate runs before the model. Ask an employee for something off its job — a recipe, a horoscope — and it refuses without ever calling the model.
Figures bound to evidence
For legal, finance, and security work, an output gate binds every number in an answer to a retrieved source. A figure it can’t bind is refused, not published.
500+ behavior scenarios
Memory under contradiction, adversarial security, tenant isolation, per-role behavior — run against the roster with weighted assertions. This is coverage of how they behave, not a field-accuracy percentage, and we won’t present it as one.
Our own observability
Errors and events are stripped of secrets and written to our own tenant-isolated Postgres — no third-party analytics SDK touches your data.
When we have a measured accuracy number worth trusting, we'll publish it with its method — the way our virality engine already refuses to show a score until it has the data. Until then, we show you the mechanism.
It stays yours.
Bring your own keys
Use your own Anthropic or OpenAI key — stored encrypted, resolved per request, billed to you.
Encrypted at rest
OAuth tokens, API keys, and integration credentials sealed with per-user envelope encryption.
No training on your data
We process your data to do the work — we don’t use it to train models. It’s our policy, in writing.
Metered honestly
Every interaction deducts Neural Credits through an atomic, audited transaction.
Self-hosted observability
Errors and events are PII-scrubbed and written to our own tenant-isolated Postgres — no third-party SDK.
No certification theater
SOC 2 is on our roadmap. We won’t claim a certification we don’t hold.
See it work for you.
Try the team free for 7 days. You set the autonomy, you approve the work.