How we build
Process questions
Why do you insist on eval harnesses?
Prompts and models change, and without evals every change is a vibe check on production. An eval suite uses real cases from your domain, scored automatically. It turns "did the update break anything" into a test run, not a support-ticket surprise. One practice most separates AI systems that survive their first year from demos that got deployed, and this is it.
What does "human review gate" mean concretely?
A named person approves AI output before it commits, whether that means sending, publishing, filing, or paying. The gate goes wherever the blast radius justifies it. Concretely, that means queues with approve, edit, or reject rather than automatic sends, plus audit logs of what was approved by whom. Gate placement is decided by consequence, not convenience. Gates can loosen later with evidence, and launching loose and tightening after an incident is the industry's standard regret.
How do you handle our data?
We use the minimum necessary, in your accounts wherever possible. Models are called with the least context that does the job. Secrets stay in your vault, not on our laptops, and we never train on your data. Provider data-handling settings are configured deliberately rather than by default. Data-flow documentation is part of every scope, because "where does our data go" deserves a diagram, not a shrug.