AI development costs: separate build, usage and operation
A practical cost model for agents, chatbots and automation, covering integrations, evaluation, platform usage and an evidence-based pilot.
An AI estimate should not start with a model name or the number of chat screens. Document questions, request classification and multi-system agents have different engineering and review needs. Separate one-time delivery, usage and operation so proposals can be compared without a headline package price hiding important limits.
Separate three cost categories
| Category | What changes it | Evidence to request |
|---|---|---|
| Build | Sources, tools, permissions and integrations | Scope, assumptions, evaluations and deliverables |
| Usage | Task volume, context size, models and steps | Estimate from representative workloads with a budget cap |
| Operation | Source updates, API changes and failures | Monitoring, alert owners and evaluation schedule |
1. Scope the actual task
For a chatbot, define questions, knowledge and unsupported topics. For automation, identify triggers, business rules and human checks. For an agent, identify permitted tools and actions requiring approval. “Do everything” cannot be estimated or accepted. Start with one important workflow and define when the system must stop because information is insufficient.
2. Data readiness carries real cost
Conflicting document versions, outdated web pages and ownerless information require review before integration. Agree authoritative sources, access rights and refresh frequency. Separate cleaning from implementation. Knowledge limited by department or account also requires permission-aware retrieval and tests; a restrictive prompt alone cannot enforce access.
3. Estimate usage from your workload
Sample short, typical and long real tasks. Measure input, output, model/tool calls and retries. One agent request may include several steps and cost more than its visible message count suggests. Run a pilot and model low, expected and high workloads using business-supplied volumes. Verify current provider prices when quoting and date the assumptions.
- Use business task volumes rather than invented forecasts
- Separate model, embedding/search, storage and connected-tool charges
- Set per-task limits, timeouts and budget notifications
- Include human review and correction work in the operating model
4. Include evaluation and control
Build test cases for ordinary tasks, missing data and unauthorised requests, with acceptable outputs or actions. Check source grounding, abstention and escalation. Removing evaluation may reduce a quote but leaves a demo without a dependable acceptance standard. A usable deployment needs both positive examples and failure boundaries.
| Use case | First testable scope | Controls |
|---|---|---|
| Knowledge chatbot | Approved small document set | Version ownership and abstention |
| Automation | Classify requests before human-approved routing | Repeated triggers, missing input and queueing |
| Agent | Draft information using limited CRM tools | API permissions, write approval and audit history |
5. Prepare an estimable brief
- Anonymised examples of real work and the expected result
- Systems, API documentation, account owners and data boundaries
- Acceptable quality, time and operating cost
- Review owners and actions that require approval
- Fallback, shutdown and correction procedures
Use the pilot to decide whether to expand
Compare the old workflow with the pilot including review time, errors and actual provider charges. Passing a test set is not enough if the team still repeats every task. Adjust the scope or use deterministic automation for appropriate steps. Expansion should follow evidence that the workflow improves useful work.
Compare the scope of AI automation →
Plan an AI pilot with a consulting scope →
