How much does it really cost to put AI into production?
We start from an explicit limit: 50,000 euros for the build and the next two years of operation. A 2-year TCO means the initial implementation plus two years of operation integrated into the company, but without the costs on the client's side. It includes software development, data integration, the model bill (e.g. Gemini), cloud costs, monitoring, and maintenance.
The short answer: 50,000 euros is enough for well-scoped AI projects. But the difference between 1,000 and 100,000 agent executions per month can change the cost completely. At the end, we discuss how you can manage success if you reach spectacular scale.
Note: We use a cap of 50,000 euros for the total cost of ownership (TCO) over a two-year period. For international readers, this threshold corresponds to a project of 50,000-60,000 USD.
1. What AI Application Can You Build Under 50K Euros
Here are a few realistic project types, with the caveat that these amounts are market estimates for low- and medium-complexity projects and do not constitute an offer.
| Project | TCO - 2 years (Low-medium complexity estimate) |
|---|---|
| AI for documents and knowledge search across company documents | 15,000-30,000 euros |
| Analysis of sales and support phone conversations, including live | 20,000-40,000 euros |
| Price & Stock Intelligence for market comparison and positioning | 25,000-45,000 euros |
| Conversational reporting over ERP/CRM for investigating your own data without predefined BI reports | 25,000-45,000 euros |
| AI agents over ERP for sales, replenishment, or pricing | 30,000-50,000 euros |
2. The Formula: One-Time Setup + 2 Years of Run
In Chapter 5 of the AI in B2B Sales Guide 2026 we split the AI budget into two parts:
Setup = architecture, software development, data integration and cleaning, business rules, testing, and launch.
Run = cloud, AI, monitoring, maintenance, support, and adjustments.
For the projects analyzed in the guide, the indicative estimate for the Run stage is approx. 5-15% of the Setup cost per year for cloud, 15-25% for maintenance and support, and 0-15% for ongoing data cleaning.
Resulting price ranges:
| Initial Setup | Indicative 2-year TCO* |
|---|---|
| 15,000 EUR | 21,000-31,500 euros |
| 20,000 EUR | 28,000-42,000 euros |
| 25,000 EUR | 35,000-52,500 euros |
| 30,000 EUR | 42,000-63,000 euros |
Note: The indicative TCO does not include additional costs on the client's side, for example on-prem infrastructure or the labor cost of your own employees. The calculation is based on the Setup/Run methodology from Guide #1 and does not constitute an offer (Guide #1)
To land in the first rows of the table (TCO under 50,000 euros), the project needs a clear objective, few integrations, and a manageable volume.
Case Studies
In an OPTI project for a distributor with more than 20,000 SKUs, AI integration with Entersoft ERP reduced the average quoting time from 23 to 7 minutes.
See the AI + Entersoft ERP case study
In a price monitoring project, an application built with BigQuery and Gemini was tracking approximately 2,000 products from 20 sources after two months.
3. How Much Does It Cost to Run AI Agents Over Systems Like ERP?
Companies want to implement AI to streamline operations running in classic software such as ERP, CRM, and WMS. But a common misconception says the whole cost is tokens. In our experience, the cost of an AI agent is not the cost of tokens.
A typical enterprise AI architecture involves the stages below, each with its own costs:
1. ERP / CRM / WMS 2. Data layer 3. Agent: tools and actions 4. Agent: AI model 5. Agent: observability 6. Action control / Human-in-the-Loop.
In short, existing applications must send unified data, and only at that level can the AI agent run. The agent is not just the model: it needs tools and must be measured and supervised, including through a final layer where human approval can come in.
The layers have different ways of calculating cost, as we documented in August 2026:
| Layer | Details | How it scales |
|---|---|---|
| ERP + Data layer | Integration, synchronization, data cleaning | mostly stable: Setup + maintenance |
| Agent harness | Orchestration, runtime, sessions, memory | Run varies with volume: per execution / runtime / storage |
| Tools | Grounding (search), OCR, APIs, email, etc. | Run varies with volume: per search / call / email |
| Model | Input, output, reasoning tokens | Run varies with volume: per token |
| Observability | Logs, traces, evaluations | Setup and Run vary with the number and complexity of executions |
| Action control / Human-in-the-Loop | Action verification, logs, and external effects | Setup mostly stable, Run varies with volume: exceptions, rules |
Google Cloud treats the AI model separately from agentic services, with model tokens billed separately. And the price difference between models is large. At current public rates for contexts under 200K:
| Vertex AI model | Input / 1M tokens | Output / 1M tokens |
|---|---|---|
| Gemini 2.5 Flash Lite | $0.10 | $0.40 |
| Gemini 2.5 Flash | $0.30 | $2.50 |
| Gemini 2.5 Pro | $1.25 | $10.00 |
Source: Google Cloud, checked September 30, 2026 (Google Cloud)
That is why an AI agent that checks whether a customer has started buying less can use a cheaper model than one that builds a complex quote.
Why Does Observability Matter?
For example, in our AI Sales 2.0 platform, we monitor and display executions, success rate, tokens, and estimated cost separately. An agent that analyzes the quality of a product listing has, in a typical demo execution:
5,400 input tokens + 920 output tokens + other tools = approx. $0.03 / run. On the demo platform, four agents monitored together show 570 runs/month, a 98.9% average success rate, and approximately $20.90 in estimated monthly cost.


Observability is not just a technical feature. For an AI workflow, Google Cloud and OPTI's detailed analysis recommend tracking the cost per successful task, not the cost per token: a cheap agent that fails often can be more expensive per useful result.
4. From 1K to 100K Runs per Month
Let's keep the example above at $0.03 / agent run. We assume the same execution complexity and the same unit cost:
| Agent runs / month | Model cost / month | Model cost / 24 months |
|---|---|---|
| 1,000 | $30 | $720 |
| 10,000 | $300 | $7,200 |
| 100,000 | $3,000 | $72,000 |
This is only the estimated AI consumption, not the full cost of the application. At 1,000 runs per month, tokens are barely relevant compared to integration and development. At 10,000, they are a small budget line. At 100,000, they exceed the cap in this article on their own.
Our previous article on AI costs in 2026 warned that the model price can fall while the application's total bill grows: agentic applications make more calls, use more services, and end up being executed more often.
Gartner estimates that an agentic task can use 5 to 30 times more tokens than a standard chatbot interaction. (Gartner)

In conclusion, the first budgeting question is the same as in software development: What process do we want to automate, how many times will it run, and what is a successful execution worth?
5. How Do I Avoid Linear Cost Growth If We Succeed?
To close the article, here are four simple recommendations for managing the linear growth of costs with volume in case of success.
Automation before AI. If a rule/trigger in the database (e.g. SQL) or a classic deterministic workflow solves the problem, don't call AI models.
Model routing. Use cheap models (e.g. Flash/Lite) for repetitive tasks, and more expensive models only when the complexity justifies it. Routing can be implemented in the application, but Google also offers Model Optimizer for automatically selecting the model tier based on cost and quality.
Smaller context and reuse cache. Don't resend thousands of ERP rows on every run (e.g. the customer's entire history). Carefully extract only the necessary context, cache wherever possible, and avoid duplicate calls. Automatic context caching services already exist for AI model calls; for example, Google enables caching by default.
Measure cost per result, not just per token. To see the linear growth factors, monitor runs, success rate, cost/run, and cost/successful task, as in the examples above. They must be observable from the application, which is why we included external observability as a mandatory layer.
Outlook: the Same AI Performance Is Getting Cheaper
An Epoch AI study published on Sept. 22, 2026 estimates that since 2023, the cost for the same level of AI performance has fallen on average by approximately 47% per quarter, or 13X / year. If the historical direction continues, well-monitored variable cost can decrease (Epoch AI).
Separately, Gartner estimates that by 2030 the inference cost borne by providers for a 1-trillion-parameter LLM will be more than 90% lower than in 2025. Gartner nevertheless warns that AI agents tend to consume as much as possible and are used more often, so the total bill can still grow. (Gartner)
In conclusion, you can build a real project with a 2-year TCO under 50,000 euros. It can analyze documents or conversations, track the market, query ERP data, or run a commercial process through AI agents. At low volumes, integration and software cost much more than tokens. At high volumes, the cost per agent run becomes critical, which requires continuous monitoring.
Have an AI project to size? Answer 5 questions:
- What process do we want to automate?
- How many systems does it need to integrate with?
- How many executions do we expect per month?
- What happens if the AI agent makes a mistake?
- What is the economic value of a successful execution?
OPTI designs and implements AI applications integrated with ERP, CRM, and company data. We work with clients in Romania and in international markets.
Resources:
OPTI Guide #1, Ch. 5 - AI Costs and Governance: Setup vs Run
OPTI - AI Costs in the First Half of 2026
OPTI - AI Agents Over On-Prem ERP
OPTI - Case Study: AI Upgrade for Sales from Entersoft
Google Cloud - Gemini / Agent Platform Pricing
Epoch AI - The Plunging Price of Thought, 22.09.2026
Gartner - GenAI inference cost forecast, 25.03.2026