MarketStarter / System 11 · Support Concierge / What it can answer

What customer service automation can’t answer.

Nearly everything written about automated customer service is written by a company selling it, which is why the useful half is missing. Here is that half: the questions it genuinely resolves, the ones it should refuse by contract, the metric the industry leads with that can be raised without improving anything, and the message volume below which none of it pays for itself.

Last reviewed Prices in EUR, excl. VAT 11 min read Διαβάστε το στα ελληνικά

The short answer

Customer service automation reliably answers four questions: where is my order, is it available in another size or colour, can I exchange or return it, and when will it arrive. In an online shop those four are typically most of the inbox — and the first one alone is around three in every five messages.

It should refuse four more, by rule rather than by judgement: money, damage, complaints, and anything the shop’s own content does not actually say. Accuracy is bounded by the source, not by the model.

Two numbers decide whether to buy it at all. Below roughly 500 conversations a month it does not pay for itself in labour. And containment rate is the wrong metric — it improves every time the system escalates less, so it is the one figure a supplier can raise without making anything better.

Why the category has a bad name, and mostly deserves it

Ask any shop owner who put a chatbot on their site in the last eight years and a good number will tell you they took it down. They were usually right to. The thing they installed answered from a decision tree: a menu of buttons leading to a menu of buttons, with no access to anything about the person in front of it. Ask it where your order is and every path ends at the same two sentences — check your email, or we will get back to you. The customer has already checked their email. That is why they are in the chat.

So the category’s reputation was set by software that could not do the single most common job, and a decade of marketing has been spent insisting the new version is different. Some of it is. The distinction that matters, though, has nothing to do with whether a product is described as AI.

The one distinction worth testing

Does it read the order, or does it recite the policy? A system with only your policy can state your return window. It cannot say “it left yesterday evening, it is at the sorting centre, it should reach you Tuesday” — and that sentence is most of what your customers are asking for.

Reading the order means two live connections at the moment of the question: into the shop platform for order and fulfilment status, and into the courier’s own tracking. Everything else in this article — the accuracy ceiling, the escalation rules, the cost — follows from whether those two connections exist.

What it can answer

Sort a month of an online shop’s messages by what was actually being asked, rather than by which channel they arrived on, and the shape is remarkably consistent. Below is an illustrative month — not a client’s result — showing where the answers come from, because that is the part that decides whether a given tool can serve them at all.

An illustrative month of e-shop support messages Illustrative · shape, not a result
What was askedMessagesShareWhere the answer lives
“Where is my order?”52863%Order record + courier tracking
“Do you have it in another size?”12114%Live stock
“Can I exchange or return it?”8410%The shop’s own policy
“When will it arrive?”405%Shipping rules + courier status
Everything else678%A person

Total 840; the four answerable classes account for 773. Read the fourth column, not the third — a tool that cannot reach the order record can only serve the 10% that lives in a policy document, whatever its marketing says.

Two things follow from that table, and both are commercial rather than technical. The first is that the biggest class is also the most mechanical: nobody asking where their parcel is wants a relationship, they want a fact that already exists in two systems they have no access to. The second is that the second-biggest class is not a chore at all — “do you have it in another size” is a customer standing at the checkout waiting to be told yes, and answered on Monday it is revenue that went somewhere else on Saturday.

What it can’t answer — and shouldn’t try

This is the half that gets left out, so it is worth being specific. Four classes of question should never be answered by automation, and in a well-built implementation they are refused by rule rather than left to the system’s judgement in the moment.

  • Anything about money. Refunds, incorrect charges, discounts, goodwill gestures. Not because a model cannot form a sentence about a refund, but because the decision belongs to somebody who can be held responsible for it.
  • Damage, wrong items and lost parcels. The customer is already unhappy, and the first reply sets the tone for whatever follows. This is the single most expensive place to be efficient.
  • A complaint in any form. Including the ones phrased politely. A complaint handled by software is the screenshot that ends up on social media, and it is a self-inflicted wound every time.
  • Anything your own content does not say. The most overlooked of the four, and the one that quietly determines the quality of everything else.

The accuracy ceiling is your content, not the model

A support agent built on your material can only be as correct as that material. Where the site is silent, it has nothing to answer from; where two pages disagree — and on return windows they very often do — there is no version of “better AI” that resolves the contradiction, because the contradiction is not a language problem. It is a decision nobody at the shop has made yet.

The honest behaviour in that situation is to escalate and report the contradiction, which turns the biggest risk in the category into a deliverable. Every question the system could not answer is a gap in what the shop says about itself, and a monthly list of those gaps is usually the first time an owner reads their own policies from the customer’s side.

The consequence for buyers

If your policies are not written down anywhere, writing them is the first piece of work — not an inconvenient detail to discover in the second month. Any supplier who does not raise this before starting has not thought about where the answers are going to come from.

And the handover is a feature, not a failure

The 8% in the table above is not leakage. It is the system doing the thing it was built to do — recognising the questions that need a person and routing them, with the conversation and the order already attached so nobody has to repeat themselves. A support agent that never escalates is precisely the one that gets removed six weeks later.

The metric to ignore: containment rate

Almost every vendor in this category leads with a containment rate, an automatic-resolution rate, or a “% of tickets deflected”. All three measure the same thing: the share of conversations that did not reach a human being.

Which makes it the one number in this industry that improves every time the product gets less helpful. Escalate less and it goes up. Widen what the agent is willing to attempt and it goes up. Remove the confidence threshold that hands over uncertain questions and it goes up nicely. Nothing about the customer’s experience has to improve for the figure to look better, and a supplier being measured on it has an incentive pointing the wrong way.

Three questions that are harder to game

Ask instead: what is the system contractually forbidden from answering? Can that rule be adjusted to improve the percentage? And what happens to a question it is not confident about? A supplier willing to put the escalation rules in the contract — rather than in the configuration — is making a claim you can hold them to.

For the same reason, be sceptical of a first-reply-time figure quoted on its own. Four seconds to say the wrong thing is worse than four hours to say the right one; the number is only meaningful next to what the system was allowed to attempt.

The volume below which it doesn’t pay for itself

This is the arithmetic no vendor page runs, so here it is. At roughly three minutes per message — reading it, finding the order, replying — 500 conversations a month is about 25 hours of somebody’s time. Below that, the hours you would save do not cover a system built to save them, and a well-written FAQ page, a set of saved replies and one inbox rule genuinely get most of the way.

There is, however, a second reason to buy it under that line, and it is entirely legitimate. It is just a different purchase:

  • Above ~500 conversations you are buying hours. The questions that did not need a person stop reaching one, and the ones that do arrive faster, with the order attached.
  • Below it you are buying response time. A reply at eleven on a Saturday night instead of nine on Monday morning is a sale that would otherwise have gone elsewhere. Worth money — just not the same money.

The failure mode is buying the second while believing you bought the first. A shop expecting saved hours at 200 conversations a month will look at the invoice in the third month and cancel, correctly. Deciding which of the two you are actually buying is a five-minute conversation that prevents it.

What it costs, itemised

The inputs cost cents. The service does not. Those two facts sit awkwardly together in most vendor pricing, so it is worth separating them explicitly.

What one automated conversation costs to run Tier B · vendor pages · Tier D · calculated
LayerCostNote
Platform, ~5 replies at $0.003 each≈ €0.013Published per-message rate; a separate unit from voice minutes
Knowledge base and retrieval€0Included in every tier, including the free one
Language model, same ~5 turns≈ €0.016Calculated — assumptions below
One conversation, all in≈ €0.03Platform plus model, no human time

Model assumptions, stated so they can be argued with: a small fast model at $1 / $5 per million tokens, about five turns per conversation, roughly 3,000 input tokens per turn including the retrieved material and the growing history with prompt caching, and about 120 output tokens per turn. Change the model and this line moves; nothing else here does.

Against that, here is what we charge. It is on this page rather than on the system page deliberately — prices age badly and get quoted back for years, so they live in one place with a visible review date.

System 11 · E-shop AI Support Concierge

Monthly, by conversation volume

  • Starter€190500 conversations a month
  • Growth€3401,500 conversations a month
  • Scale€4903,000 conversations a month

Build: €1,690 for one language, the order and courier lookups and the escalation path — up to €2,900 with the phone channel, a second language, returns flows or more than two couriers.
Over the tier: €0.08 per conversation. A second language: €90 per month. Adding the phone: €290 per month, which is System 01’s Starter tier and carries its own meter of minutes.
The shop platform subscription and any paid apps stay on your own bill. Prices exclude VAT.

Three cents of inputs against thirty-eight cents of Starter tier is a ratio somebody will do the arithmetic on, so let us do it first. The monthly fee is not margin on the message. It is the build amortised, and the hour every month that keeps the knowledge base honest — closing the gaps the escalations exposed, which is the work that stops the system slowly getting worse. The tokens are the cheapest thing in the stack, and they are not what you are buying.

Two more things about the pricing that are worth stating because their absence causes arguments:

  • Never priced per resolved ticket. How many people ask is set by your customers; how well the system answers is set by us. Only one of those should carry the risk, and per-resolution pricing puts it on the wrong side.
  • A metered unit needs a definition before the first invoice. Ours: one visitor session with at least one reply, inside 24 hours. Continuations after that count as new, and sessions identified as bot or spam traffic do not count at all — which is both our own cost guard and the reason the number can be trusted.

One note on the Greek market specifically: we went looking for published local pricing for this category while doing our own pricing research and found none. Every Greek quote we are aware of is issued privately on request. That is not a complaint — it is the reason this page exists in the shape it does.

Chat or the phone — the economics aren’t close

Both channels can run on the same knowledge and the same integrations, which makes the choice look like a preference. It isn’t. Speaking is the expensive part of talking: on a call, the voice layer alone costs around nineteen times what the language model behind it does. That single fact explains why a phone agent is sold in minutes and a chat agent in conversations, and why nobody should quote you one unmetered.

Same brain, two channels Tier D · calculated from published rates
ChatPhone
Metered inConversationsMinutes, and concurrent calls
Raw cost per unit≈ €0.03 per conversation≈ €0.072 per minute
Right whenVolume is written, constant, order-shapedThey want a voice now, or are driving
SharedKnowledge base · order and courier lookups · escalation rules — built once

The per-minute phone figure is itemised in full in what an AI receptionist really costs, from our own invoices.

The practical consequence: whichever you build first, the second is an add-on rather than a project, because the expensive work is the knowledge and the connections, not the channel. What does not work is being quoted for the two as unrelated builds — and if a supplier tells you the calls come free with the chat, they are selling you a surprise.

Size it against your inbox Bring one number: how many messages you answered yesterday, and how many were about an order that had already shipped. Or write to hello@marketstarter.gr.

Sources, and how confident we are in each number

The figures above are not all the same kind of claim, so they are graded rather than presented as equivalent.

  • Vendor price pages — the $0.003 text message, the knowledge base and retrieval included at every tier, the published voice rate. Public and checkable, as of August 2026.
  • Calculated — the per-conversation model cost and the €0.03 all-in. Assumptions are stated inline: model rates, turns per conversation, tokens per turn, and $1 ≈ €0.86. Change the model or the rate and these move; nothing else does.
  • Illustrative — the 840-message table. It is the shape we consistently see in e-shop inboxes, not a client’s reported result, and it is labelled that way because a percentage borrowed from somebody else’s shop tells you nothing about yours.
  • Our own research — the absence of published Greek pricing for this category, from the competitor pass behind our own price list.
  • No performance claims. There is deliberately no containment rate, no resolution percentage and no case study on this page. This system is new to our menu, and given what the fourth section says about how manufacturable those numbers are, publishing one would be self-refuting.

A word on shelf life. Model prices in particular fall quickly, and an article quoting them gets cited for years. This page carries a visible review date and is checked when the underlying rates move. If you are reading it long after , treat the cost lines as historical and ask us for the current ones.

MarketStarter is a HubSpot Solutions Partner and a Salesforce Partner, based in Greece, answering everywhere. The system this article describes is System 11.

Straight answers

The questions people actually ask.

What is customer service automation?

Customer service automation is the use of software to answer routine customer questions without a person handling each one. In an online shop, the useful version does two things rather than one: it answers from the shop’s own content — products, variants, policies, shipping rules — and it looks up the live order at the moment of the question, in the shop platform and in the courier’s own tracking. The distinction that matters commercially is not whether a system is described as AI. It is whether it reads the order or recites the policy. A system that only has the policy can tell somebody what the return window is; it cannot tell them where their parcel is, which is most of what they are asking.

What can customer service automation actually answer?

Four questions, and in an online shop they are most of the inbox. Where is my order, which needs a live lookup rather than a policy. Is it available in another size or colour, which needs current stock. Can I exchange or return it, which needs the shop’s own written policy. And when will it arrive, which needs the shipping rules and the courier’s status. In an illustrative month of 840 messages those four account for 773 of them. The first one alone is usually around three in every five.

What can customer service automation not answer?

Four classes of thing, and a good implementation refuses them by rule rather than by judgement. Anything about money — refunds, charges, discounts, goodwill. Anything about damage, a wrong item or a lost parcel. A complaint in any form. And anything the shop’s own content does not actually say. That last one is the ceiling nobody mentions: accuracy is bounded by the source, not by the model, so where two pages of a site disagree about the return window, the honest behaviour is to escalate and report the contradiction rather than pick one. If your policies are not written down anywhere, that is the first piece of work, not a detail to discover in month two.

Is a containment rate a good way to compare vendors?

No, and it is worth understanding why, because it is the number most vendors lead with. A containment or automatic-resolution rate measures the share of conversations that did not reach a person — which means it improves every time the system escalates less. It is the one metric a supplier can raise without improving anything, simply by tuning down the handover. Ask instead what the system is contractually forbidden from answering, whether that rule can be adjusted to make the percentage look better, and what happens to a question it is not confident about. A supplier who will put the escalation rules in the contract is making a claim you can actually hold them to.

How many support messages do you need before automation pays for itself?

On labour, roughly five hundred conversations a month. At about three minutes a message that is some twenty-five hours, and below it the hours saved do not cover a system built to save them — a well-written FAQ page, a set of saved replies and one inbox rule get most of the way. There is a second and entirely legitimate reason to buy below that line, but it is a different purchase: a reply at eleven on a Saturday night instead of Monday morning is a sale that would otherwise have gone elsewhere. That is response time, not saved hours. Shops that think they are buying the first and receive the second tend to cancel in the third month.

How much does customer service automation cost?

The inputs are cents and the price is not, which is worth separating. On published vendor rates a five-reply conversation costs about €0.013 of platform messages plus roughly €0.016 of language model, so the raw cost of one lands near €0.03. MarketStarter charges €1,690 to €2,900 to build, then €190 a month for 500 conversations, €340 for 1,500 or €490 for 3,000, with anything over the tier at €0.08 a conversation and a second language at €90 a month. The gap between three cents and thirty-eight is not margin on the message: it is the build, and the hour every month that keeps the knowledge base honest. Prices exclude VAT.

Will it give my customers wrong information?

It will if it is allowed to answer from a general model rather than from your material. The safeguard is architectural rather than a matter of trust: the agent answers only from the shop’s own content and from live order data, and anything below a stated confidence threshold is handed over instead of guessed at. The useful by-product is a monthly list of everything it could not answer, which is the first time most owners read their own policies from the customer’s side. Most discover that two pages of their site contradict each other about returns.

Should it be chat or the phone?

Follow the volume, and know that the economics of the two channels are not close. Speaking is the expensive part: on a call the voice layer alone costs around nineteen times what the language model behind it does, which is why a phone agent is metered in minutes and a chat agent in conversations. So chat suits volume that is written, constant and shaped like where-is-my-order, and the phone suits the caller who wants a voice now or is driving. The part worth knowing before you buy either is that the knowledge base, the integrations and the escalation rules are one layer, so whichever you build first makes the second an add-on rather than a second project.

Count yesterday’s messages first.

How many you answered, and how many were about an order that had already shipped. That one number decides all of this. 20 minutes, no obligation.