T2ETECH
ai / GUIDE

AI development costs: separate build, usage and operation

A practical cost model for agents, chatbots and automation, covering integrations, evaluation, platform usage and an evidence-based pilot.

T2ETECH Editorial3 min read

An AI estimate should not start with a model name or the number of chat screens. Document questions, request classification and multi-system agents have different engineering and review needs. Separate one-time delivery, usage and operation so proposals can be compared without a headline package price hiding important limits.

Separate three cost categories

Costs to itemise in the proposal
CategoryWhat changes itEvidence to request
BuildSources, tools, permissions and integrationsScope, assumptions, evaluations and deliverables
UsageTask volume, context size, models and stepsEstimate from representative workloads with a budget cap
OperationSource updates, API changes and failuresMonitoring, alert owners and evaluation schedule

1. Scope the actual task

For a chatbot, define questions, knowledge and unsupported topics. For automation, identify triggers, business rules and human checks. For an agent, identify permitted tools and actions requiring approval. “Do everything” cannot be estimated or accepted. Start with one important workflow and define when the system must stop because information is insufficient.

2. Data readiness carries real cost

Conflicting document versions, outdated web pages and ownerless information require review before integration. Agree authoritative sources, access rights and refresh frequency. Separate cleaning from implementation. Knowledge limited by department or account also requires permission-aware retrieval and tests; a restrictive prompt alone cannot enforce access.

3. Estimate usage from your workload

Sample short, typical and long real tasks. Measure input, output, model/tool calls and retries. One agent request may include several steps and cost more than its visible message count suggests. Run a pilot and model low, expected and high workloads using business-supplied volumes. Verify current provider prices when quoting and date the assumptions.

  • Use business task volumes rather than invented forecasts
  • Separate model, embedding/search, storage and connected-tool charges
  • Set per-task limits, timeouts and budget notifications
  • Include human review and correction work in the operating model

4. Include evaluation and control

Build test cases for ordinary tasks, missing data and unauthorised requests, with acceptable outputs or actions. Check source grounding, abstention and escalation. Removing evaluation may reduce a quote but leaves a demo without a dependable acceptance standard. A usable deployment needs both positive examples and failure boundaries.

Choose a pilot around readiness
Use caseFirst testable scopeControls
Knowledge chatbotApproved small document setVersion ownership and abstention
AutomationClassify requests before human-approved routingRepeated triggers, missing input and queueing
AgentDraft information using limited CRM toolsAPI permissions, write approval and audit history

5. Prepare an estimable brief

  • Anonymised examples of real work and the expected result
  • Systems, API documentation, account owners and data boundaries
  • Acceptable quality, time and operating cost
  • Review owners and actions that require approval
  • Fallback, shutdown and correction procedures

Use the pilot to decide whether to expand

Compare the old workflow with the pilot including review time, errors and actual provider charges. Passing a test set is not enough if the team still repeats every task. Adjust the scope or use deterministic automation for appropriate steps. Expansion should follow evidence that the workflow improves useful work.

References

OpenAI — Evaluation best practices