PILOT
Start with a workflow.
For a team proving a use case with governed workspace access or a shared inference endpoint.
- Focused onboarding and model shortlist
- Usage and quality review
- Clear path to production
PRICING / START WITH THE WORKLOAD
The final Dvar rate card is still being shaped with early enterprise partners. The structure is clear now: start with an eval, pay for the traffic you run, and only commit to capacity when the workload proves it.
ILLUSTRATIVE STRUCTURE
These are placeholders for the first public rate card. We will publish concrete model and compute rates after validating demand, regional costs, and support requirements with design partners. Support is part of the proposal, not a surprise line after go-live.
A quality, latency, and unit-cost baseline built with your team from representative traffic.
Usage-based routing to the cheapest model that clears your eval, with serving and cost visibility.
Evaluation, data curation, training, deployment, and an operated improvement loop when volume earns it.
PILOT
For a team proving a use case with governed workspace access or a shared inference endpoint.
PRODUCTION
For a product team that needs predictable capacity, policy controls, and a production operating rhythm.
ENTERPRISE
For a strategic workflow where data, evaluation, and specialized models create a durable quality wedge.
THE HONEST ANSWER
Model API price is only one line on the bill. A useful decision includes tokens, latency, GPU occupancy, support, data boundaries, and the cost of a bad answer. That is what the workload review is for, and our team stays accountable for the path after the first deployment.
Ask for the rate card when it is ready