Pricing
Your first 20 tests are free.
Preview agent-specific tests with no card. Open the full validation dataset for $10 one time, per agent. Hand us the build when it is time to ship.
Faultmap Preview
Agent-specific validation tests generated from the context you provide.
Preview my tests- 20 agent-specific validation tests
- Generated from the goal, users, data, and tools
- Shareable validation preview
- No card, no code
Production plan
Three production agents, built and kept cost-efficient.
Start free- Three production agents built to pass their own Faultmap
- Kept cost-efficient continuously, savings baselined from day one
- Exhaustive pre-ship validation on every path
- Renewal is a report, not a pitch
- $3K per additional agent
Full validation dataset
The complete validation dataset generated for your agent and use case.
Get the full dataset- Complete dataset for the context you provide
- Coverage across goal, users, data, and tools
- Ranked by what breaks first
- Export and share with your team
Fixmap
The fix plan: the architecture to change and the tests to pass.
Get the fix plan- Everything in the full validation dataset
- The architecture to change, and why
- Tests to pass, ranked by what breaks first
- Build it right the first time, or hand it to us
Forge
A production agent, built to pass its own Faultmap.
Talk to us- Production agent built for you
- Passes its own Faultmap
- Monitoring included
- Two partners signed pre-launch
Self-serve and card-friendly. No sales call to preview your first 20 tests.
The $10K/year plan, anchored
vs $55–110K of build labour for one agent
vs $150–600+ per 1K requests on unvalidated frontier paths
vs $20–100K a year fixing what breaks in production
Renewal is a report, not a pitch: savings baselined from day one, quoted before you commit.
Cost figures are approximate and vary by workload, model, region, and provider terms.
Questions
Before you start. Answered.
The pre-build AI agent doctor. We map where your agent will break before you build it, prescribe the fix, and forge the agent that passes.
A failure map of where your agent breaks, plus the first test suite it has to pass, generated from your goal and data before any code.
No. It is built for the pre-build moment, where a fix costs cents instead of thousands. The same failure-pattern library also diagnoses and fixes agents already breaking in production.
They start after you build. We start before. We map where the agent breaks before a line of code exists.
It runs on your side in the design phase. No code, prompts, or traces leave your environment.
The library is built from real agent behaviour across production builds and public failure corpora: retry spirals, silent successes, tool-permission blast radius, cost runaway, and every class in our failure catalogue. Each pattern is a concrete failure signature with its root cause, pre-build signal, and a test that catches it. We publish the taxonomy and the counting method in the Hard 70% series; the honest number is "five million and growing" — the moat is the classes, not the count.
Routing picks a cheaper model at runtime. It does not prove the task still succeeds. We prove it before the build: the failure map, the path enumeration, and the test suite exist before code does, and the cost quote is generated from the same map. Routing is one technique inside a right-sized architecture; the proof is the product.
For tasks where the failure library shows a leaner model and architecture are sufficient — which is most production agent steps — yes, and we show the band, not a cherry-picked point: 93–97% inside the frontier's 90–97% agentic band, at roughly 20–200×+ lower run cost. Tasks that genuinely need frontier reasoning get routed there, and the quote says so before you commit. Every agent ships with its own cost justification, so renewal is a report, not a pitch.
Because the alternative is six figures: $55–110K of build labour, $150–600+ per 1K requests on unvalidated frontier paths, and $20–100K a year fixing what breaks in production — for one agent. $10K buys three production agents built and kept cost-efficient, with savings baselined from day one.
Still deciding? Preview your first 20 tests.
No card. No code. See the validation data before you pay.