I'm building 1trAIner — an AI coach that creates plans, reviews workouts, and answers in chat. It seems easy to calculate the cost of such a service: look at model charges, divide by users, and get the cost per person per month.

But the bill includes more than just conversations with the coach. There are test responses, translations, editorial proofreading, and re-evaluations after corrections. All of these are model costs, but their reasons differ. If you add them up without breaking them down, you get a number that's inconvenient for making decisions.

Let me set a boundary right away: this article will not include the monthly DeepSeek bill or the cost of serving one user. The service is young, the audience is small, and we don't yet have an honest monthly sample. I don't want to substitute it with expenses for individual tasks.

However, there are specific charges, the structure of our calculation, and a story about why work stopped while the budget was unspent. These examples clearly show what exactly needs to be counted before writing "our AI costs so much per month".

Not all model calls are user service

Inside the service, DeepSeek acts as a coach. But the same model can translate a service message or participate in checking a new rule. The model name is the same, but the purpose of the expense is not.

For myself, I separate the coach's daily work from product changes. In the first case, a person asks a question, gets a review or a plan. In the second, we translate the interface, check responses, and fix wording. I look separately at the costs of the editorial model: it doesn't conduct a dialogue with the user but evaluates or proofreads what we plan to show them.

This is not an attempt to remove inconvenient expenses from the calculation. They don't disappear. The questions are just different:

  • How much does it cost to serve the current load?
  • How much did we spend on changing the product?
  • How much did checking that change cost?

For example, when checking the medical boundary in the coach's responses, we received 36 test responses from DeepSeek. They went through the live service and ended up in operations, not in the task budget. Their cost is on the order of cents.

You can't get an exact cost per response from such an estimate. Even less can you stretch it over a month. All you can say is that test requests also consume money and that the boundary between development and operations doesn't appear in the bill on its own.

Even the cost of a request can be calculated incorrectly

In our calculation, it's not enough to take all the text sent and multiply its volume by the price.

The model receives input text and returns a response. Part of the input may repeat and be billed at a separate rate: the provider counts it as already processed. An important detail is that this reused volume is already included in the total input counter.

If you first pay for the entire input and then add the reused part, you'll count it twice. That's why the calculation first separates regular input from reused input, and then applies the corresponding rates.

There's a similar subtlety with the model's internal reasoning. In our calculation, it's already included in the output volume and isn't added separately. Otherwise, your own expense table would show more than the source counters imply.

To the reader, this may be boring accounting. To me, it's the foundation of any conversation about savings. Before cutting requests or switching models, you need to make sure we haven't invented extra costs ourselves.

And it's also important not to overestimate such a calculation. Having a correct function doesn't prove the monthly report is complete. It answers the question "how to calculate a specific request," but it doesn't confirm that all requests made it into the final sample.

Today's price list shouldn't rewrite last month's bill

Our price is selected based on the supplier, the model name, and the time of the request. The calculation looks for the suitable plan that was already in effect at that time.

Why do I even pay attention to this? Because the temptation to recalculate the entire history at the current price list is very strong. You get a neat table, but it answers a different question: how much the previous load would have cost under today's conditions.

For the actual expense, you need the price at the time of the request.

The calculation also includes standard and peak rates. This is not a claim that every request we make necessarily becomes more expensive at a particular time. It is an option to apply such a rate if one is set. If a separate peak rate is missing, the standard one is used.

I wouldn't turn this into advice to immediately move all work to cheaper hours. A chat reply and background proofreading have different urgency. First we need to see which workload actually falls under a different rate and what can be postponed without harming the scenario.

The cost of a decision here is not only monetary. You can save on a request but make a person wait where waiting breaks the point of the product. Without load data, such a trade-off is easy to make blindly.

Specific amounts: proofreading, not a month of coach work

The most noticeable expenses of recent days are proofreading English and Chinese pages. This is the work of an editorial model via OpenRouter, not a bill for DeepSeek's everyday replies.

In the initial proofreading stage, $3.89 was spent. In the continuation after topping up the balance — $4.98. In the evening stage — $4.71. These are expenses of individual stages, not the service's rate or the cost of a month.

Behind them are quite concrete corrections. In the English translation, the marathon name "Дёмино" turned into the invented "Demyansky". The editor restored the correct name. In the help section, it clarified the wording about Garmin data: it refers to recorded data, not to the watch supposedly being the only source of any facts.

At the same time, we don't trust the editor unconditionally either. Edits go through automatic checks: whether numbers are preserved, whether the terminology glossary is followed, whether there is any unexpected mixing of alphabets. If a check rejects an edit, it doesn't make it into the published translation.

There is an unpleasant side to saving here. The money for the model's work has already been spent, but some suggestions are not used. One could consider this a loss and weaken the checks. That's not how I look at these expenses: a paid reply doesn't become correct just because it was billed.

Keeping the previous machine translation also doesn't mean getting perfect text. It only means not applying a change that failed the current rules. The question of quality remains open and requires a separate decision.

There's a budget, but work is stalled

During proofreading, a situation arose several times that is easy to misdescribe with the words "we ran out of money".

A budget was allocated for the task, and it had not yet been spent. But OpenRouter refused the next request: the available balance was not enough for a request with the specified maximum response size.

The spending limit for a task and the provider's available balance are different things. The limit sets how much we allowed to spend. The balance determines whether the provider can accept the next request with its parameters.

In the initial stage, $3.89 was spent against a budget of up to $10, but the continuation still stopped. After topping up, the review continued. Later, the balance restriction arose again.

There was also a less obvious case: the editor had already returned edits, they were applied, and the subsequent request for a summary no longer went through. So an error message at the end did not mean that the entire stage turned out to be empty.

From this I derive a practical accounting rule: look not only at the status "completed" or "error", but also at the result actually obtained. Otherwise you might either fail to account for work done or pay twice for what has already been completed.

In these tasks, no retries were launched after a refusal due to balance. Continuation was tied to saved results. This is not a fancy model cost optimization, but ordinary care: not making the system do again what it has already done.

Quality checks also have their own bill

A separate story is the boundaries of the coach's answers. When checking language versions, it turned out that the problem was not only in translation: answers in Russian also required corrections.

In the task, the model instructions were changed, automatic checking was added, and the answers were re-evaluated. The cost was $1.42. Within this amount, translating the warning via DeepSeek cost $0.01, and the rest in the table was editorial evaluation runs. The coach's own test answers, as I already wrote, were accounted for separately.

It is important that the spent budget did not mean achieved quality: we never met the target criterion. The presence of a warning at the end of the answer did not always fix an unsuccessful wording at the beginning.

In the next stage, a more detailed automatic check was added. The evaluation cost $0.41. The target Chinese answer received a higher score, but the average score for the same set of answers dropped from 4.33 to 4.00. The editor raised some objections to words that had previously been rated more leniently.

For me this is more useful than the story "we spent a little and fixed everything". Several limitations are visible at once: an added line does not necessarily eliminate a contradiction, and a model-based evaluator is not always stable.

So an editorial score is not proof of quality for me. We now check the required features of an answer against a fixed list, and use the editorial model to revise the wording.

What it takes for an honest monthly price

After that breakdown, showing the provider's remaining balance is not enough for me. A monthly cost requires expenses for a specific period, clear rules for which requests are included, and user activity data for the same period.

You need to decide in advance what counts as operations, how to account for test requests on a live service, and where to attribute the translation of new pages. And not change these rules from report to report just to get a nicer result.

I also wouldn't divide the cost by registered accounts without caveats. A person who never opened the coach and a person who regularly talks to it create different loads. An average can be useful, but only if it's clear what exactly is in its denominator.

So my conclusion is more modest than "an AI coach costs pennies." In the tasks reviewed, the costs of model requests are small. There is a separately paid quality check. There are stops due to balance even though the task budget is not exhausted. And there are results that cannot be considered sufficient based solely on the amount spent.

The monthly cost starts with accounting for these differences. And the price of the whole service does not end with the DeepSeek bill.