AI commerce agent buyer checklist: 12 questions before you buy
Twelve questions to ask any AI commerce agent vendor, with good signals, warning signs, and a scoring sheet you can copy — usable against bitbybit and its competitors.
Most AI commerce agent evaluations are decided by a demo, and a demo is a controlled environment built by the people being evaluated. It tells you the system’s best case. It tells you almost nothing about how it behaves on your catalogue, with your policy exceptions, in month four, when the person who set it up has moved on.
The twelve questions below are designed to close that gap. They are written to be used against any vendor in this category — including bitbybit — and several of them are questions we would not score full marks on for every buyer. That is the point: a checklist where every criterion resolves in one vendor’s favour is a sales document, not an evaluation.
How to use this checklist
Nine categories, twelve questions. Score each 1–5 against the good signals and warning signs, then weight the ones your business genuinely depends on rather than averaging everything equally — a merchant whose whole operation runs on one commerce platform should weight the commerce and portability questions far above the rest.
Two rules make the scores mean something:
- Require demonstration, not description. “Yes, we support that” is a claim. “Here it is, on your data, and here is the log of what it did” is evidence. For anything you score 4 or 5, you should be able to say what you saw.
- Ask the vendor to name a poor-fit case. Any honest vendor can describe a customer they would turn away. An inability to do so is itself a finding — it usually means the answer to every question will be yes.
| Category | What it tests | Questions |
|---|---|---|
| Capability | What the system can actually do | 1 |
| Context | Whether it knows who it is talking to | 2, 5 |
| Commerce | Whether it can transact, and against which source of truth | 3 |
| Control | What constrains its behaviour and actions | 6, 10 |
| Quality | How you know it is answering correctly | 8 |
| Security | Data handling, access, and review posture | 9 |
| Observability | Whether you can reconstruct what happened | 7 |
| Portability | Whether you can leave | 4, 11 |
| Economics | What it costs at your real volume | 12 |
The 12 questions
1. What can the agent actually do — and which of those things change state?
Ask: For the complete, enumerated list of actions the agent can take. Then ask which of them are read-only lookups and which change something in a real system. Then ask the single most consequential thing it can do without a human involved.
Good signal: A finite, named list, with read and write actions distinguished, and a straight answer to the “most consequential action” question.
Warning sign: “It can do anything, it’s an AI” — or “it can call our API,” which is an unbounded permission described as a feature. Also: an inability to distinguish between answering a question about an order and changing one.
2. What business systems does it connect to, and how deep does each connection go?
Ask: Not whether it “integrates with” your commerce platform, helpdesk, or ads account, but what each integration actually reads and writes. Live inventory or a nightly sync? Order lookup only, or order creation? Which direction does customer data flow?
Good signal: Per-integration specificity, and a documented list you can read before the call. A vendor that volunteers what an integration does not cover is being straight with you.
Warning sign: A logo wall where every connection is described identically. “Custom integration available” as the answer for the system you actually depend on — that is a services engagement wearing a product’s clothes.
3. Which system is the source of truth for catalogue, inventory, and orders?
Ask: Where price, stock, and orders live once this is running, and what happens when the two systems disagree.
Good signal: One unambiguous answer. Either your existing commerce platform remains authoritative and the agent reads from it and writes orders back, or the vendor supplies the commerce backbone and is honest that it is now the system of record. Both are legitimate.
Warning sign: Vagueness, or an architecture where both systems hold inventory. Your customers will find the disagreement before you do. The architecture guide covers why this is the load-bearing question in the commerce layer.
4. Who owns the customer data, and what does the vendor retain?
Ask: Whether your customer records, conversation history, and tags are your data; what the vendor retains after termination and for how long; and whether your data is used to train shared models.
Good signal: A clear ownership statement, a documented retention position, and a straight answer on model training. The answer should be in writing, in the contract or a published policy — not verbal.
Warning sign: Ownership discussed only in terms of access (“you can always see your data” is not ownership). Evasion on post-termination retention. A training answer that requires a follow-up question to get.
5. Can you export everything, yourself, without asking?
Ask: For a full export — records, conversation history, tags, and consent state — during the trial, performed by you, not by their support team.
Good signal: A self-serve export you actually run and inspect before signing. Check that it includes conversations and tags, not just a contact list.
Warning sign: Export by support request. Export that returns contacts but not conversation history. Any framing where portability is a retention-conversation topic rather than a product feature. This question is cheap to ask and expensive to skip: it is the one that determines whether the other eleven are reversible decisions.
6. How is the agent grounded in your business knowledge?
Ask: Where the agent gets the answer to “is this in stock,” “what is your return policy,” and “where is my order.” Specifically: is it retrieved when the question is asked, or was it written into a prompt during setup?
Good signal: Volatile facts — price, stock, order status — retrieved per answer from the system that owns them. Policy content authored in one place and retrieved. A clear story for how the knowledge stays current when someone changes a policy.
Warning sign: Setup that consists of pasting your catalogue and policies into a text box. It demos beautifully and goes stale the first time a price changes. Test this directly: change something in your catalogue, then ask the agent about it.
7. What happens when the agent is uncertain?
Ask: What the system does when it cannot find a grounded answer. Then test it — ask something your knowledge base does not cover and see what comes back.
Good signal: Uncertainty is a routing decision. Low confidence, a missing source, or an out-of-scope request deterministically produces a clarifying question or a handoff, and you can see that rule.
Warning sign: A fluent, plausible, wrong answer. In commerce a confident wrong answer costs more than a slow one, because it produces a failed order and a support case rather than a delay. Also a warning sign: confidence scores that appear in a dashboard but change nothing about behaviour.
8. How does human escalation work, end to end?
Ask: Which conditions force a handoff, whether you can define your own, what context the person receives, whether the customer has to repeat anything, and what happens after the person resolves it.
Good signal: Explicit, editable triggers; full conversation and action context delivered to the human; a defined return path that updates the customer record.
Warning sign: Escalation as an unhandled exception — the agent stops and someone notices. A handoff where the human opens a separate tool and asks the customer to start again. Designing clean escalation covers what good looks like here in more detail.
9. How do you evaluate answer and action quality over time?
Ask: How the vendor knows quality has not degraded after a prompt change, a knowledge update, or a model upgrade — and what you can run yourself.
Good signal: A fixed evaluation set that is re-run on change, plus the ability to review and label real conversations. Analytics counts what happened; evaluation scores whether it was right, and you want both.
Warning sign: Deflection rate offered as the quality metric. Deflection measures whether a human got involved, not whether the customer was served — a system that stops recognising when it is stuck looks excellent by that number. The reasoning is set out in measuring AI agent quality.
10. What observability and action history exists?
Ask: To reconstruct, for a specific past conversation, what the agent said, what it retrieved to justify that, what actions it took in which systems, and where it escalated.
Good signal: You can answer all four yourself, in the product, without a developer or a support ticket.
Warning sign: Prompt and response logs presented as observability. They tell you what was said, not what was done — which is precisely the gap you will need closed the first time an order is created incorrectly.
11. How are permissions and tools controlled?
Ask: Who can change what the agent is allowed to do, whether changes are logged, what team roles exist, and whether you can restrict an action to specific conditions.
Good signal: Per-action permissions, workspace roles, and an audit trail on configuration changes as well as on agent actions.
Warning sign: A single admin surface where anyone with access can widen the agent’s authority silently. Ask how you would discover that someone had.
12. What does deployment actually require, and what is the real total cost?
Ask: For a written worked example at your real volume. Then ask what the first four weeks require from your team, in hours and in whose hours.
Good signal: A cost model with the parts separated — subscription at the tier that includes what you need, any usage charges, the messaging channel’s own fees, and onboarding. On WhatsApp, note that the channel bills separately: Meta charges per delivered message and rates vary by message category and country, so any comparison between vendors has to hold that constant. A vendor who quotes it as pass-through at published rates is making a checkable claim; ask what markup, if any, sits on top.
Warning sign: A single monthly number with no volume sensitivity. Capabilities demoed on an enterprise tier and priced from the entry tier. Onboarding described as “we’ll handle it” without an hours estimate for your side.
Buyer red flags
Distinct from a low score on one question, these are patterns that should make you slow the whole process down.
- A polished demo and no evaluation framework. If the vendor cannot describe how they know quality is holding, the demo is the only evidence that exists.
- Data ownership answered in terms of access. “You can always see your data” is not the same as owning it or being able to take it.
- Human handoff that is manual or undefined. Usually means it was added after launch, which means context does not travel.
- No export path you can run yourself. Makes every other decision irreversible.
- Usage charges that surface only at scale. Ask for the price at three times your current volume, in writing.
- An “AI agent” that is a scripted FAQ bot. Test it with a question phrased in a way nobody would have anticipated. If it deflects, it is a chatbot with better copy.
- No way to constrain business actions. If the agent’s permissions cannot be narrowed, the blast radius is whatever the integration allows.
- No observable action history. You will find out you needed this at the worst possible moment.
- Integration that depends on bespoke work. Custom work built for you is custom work maintained for you, and it breaks when either side updates.
- A vendor who cannot name a customer they would turn away. Every real product has a poor-fit case.
When an AI commerce agent is the wrong purchase
Three situations where the honest answer is to wait, regardless of vendor.
- Your catalogue and policies do not live in a system anything can read. Grounding is the foundation. If price and stock live in a spreadsheet someone updates by hand, you will spend more time maintaining the agent’s knowledge than it saves. Fix that first.
- Your volume is genuinely low. If one person answers everything within a reasonable time, the remaining value is coverage outside working hours and follow-up that would not otherwise happen. That can still be worth it — but it is a smaller case than the one you will be sold, and it should be priced accordingly.
- Nobody will own it. These systems need a person responsible for keeping the knowledge correct and reviewing what the agent handled badly. That is a real, ongoing role, and a deployment without an owner degrades quietly.
The scoring sheet
Copy this into your evaluation doc. Score 1–5, weight the rows that match your dependencies, and record what you actually saw rather than what you were told.
AI COMMERCE AGENT EVALUATION — <vendor> Date: ________
Evaluator: ____________ Trial on our own data: Y / N
CAT # CRITERION SCORE WEIGHT EVIDENCE SEEN
CAP 1 Enumerated actions; read vs write clear __/5 __ ______________
CTX 2 Integration depth, per system __/5 __ ______________
COM 3 Single, stated source of truth __/5 __ ______________
POR 4 Data ownership + retention in writing __/5 __ ______________
CTX 5 Self-serve full export (records+convos+tags) __/5 __ ______________
CTL 6 Grounding retrieved, not pasted __/5 __ ______________
CTL 7 Uncertainty routes (tested, not described) __/5 __ ______________
QUA 8 Escalation: triggers, context, return path __/5 __ ______________
QUA 9 Evaluation set re-run on change __/5 __ ______________
OBS 10 Action history reconstructable by us __/5 __ ______________
SEC 11 Per-action permissions + config audit trail __/5 __ ______________
ECO 12 Written cost model at 1x and 3x our volume __/5 __ ______________
RED FLAGS OBSERVED: ________________________________________________
POOR-FIT CASE THE VENDOR NAMED: ____________________________________
DECISION: proceed / trial extended / decline Reason: ____________
Two notes on scoring. A single 1 on questions 4, 5, or 7 should stop the process regardless of the total — those are the ones that are either irreversible or actively dangerous. And do not let a strong score on question 1 carry the sheet; capability is the easiest thing to demonstrate and the least predictive of how the system behaves in production.
Where to go next
If you want the conceptual model these questions come from, the AI commerce platform architecture guide sets out the eight layers each question probes. For the security questions specifically, is the WhatsApp Business API secure? separates transport encryption from platform storage — a distinction vendors routinely blur — and our own answers on data handling, access management, and incident response are on the security and trust page. Common questions about how bitbybit works, including pricing and data, are answered on the FAQ.
Use this checklist when evaluating any AI commerce platform — including bitbybit.
Frequently asked questions
What should I ask an AI agent vendor before buying?
Start with the four highest-signal questions: the complete list of actions the agent can take and which of them change state in your systems; who owns the customer data and whether you can export all of it yourself; what the system does when it is uncertain, and whether that is an enforced routing rule or a hope; and what the total cost is at your real volume including the messaging channel's own per-message fees. Every other question is easier to recover from getting wrong.
How do I evaluate an AI agent beyond the demo?
Ask for a trial on your own catalogue and your own policy documents, then test three things a demo never covers: a question whose answer changed recently, to see whether grounding is retrieved or baked in; a request the agent should refuse, to see whether it escalates or improvises; and a returning-customer scenario in a new conversation, to see whether the system recognises them. Then ask to see the action history for those three conversations.
What are the warning signs when buying an AI commerce agent?
A polished demo with no evaluation framework behind it; an unclear or evasive answer on who owns the data; human handoff that is manual or undefined; no self-serve export path; usage charges that only appear at scale; an 'AI agent' that turns out to be a scripted FAQ bot; no way to constrain what business actions the agent may take; no observable action history; and integrations that depend on bespoke work the vendor will have to maintain for you.
Who owns the customer data in an AI commerce platform?
That depends entirely on the contract and the product, which is why it must be asked explicitly rather than assumed. The answer worth having has three parts: your customer records, conversations, and tags remain your data; you can export them in full yourself without raising a request; and the vendor states what it retains after termination and for how long. A vendor that can only answer the first part has answered the easy third of the question.
How should human escalation work in an AI commerce agent?
It should be a designed path with explicit triggers, not a fallback that fires when the agent gets stuck. Ask which conditions force a handoff, whether you can add your own, what context travels to the person picking up, whether the customer has to repeat anything, and what happens after resolution — does the agent resume, and does the customer record get updated. A handoff that loses context is worse than no automation.
How much does an AI commerce agent actually cost?
The platform subscription is usually the smaller half. Model the full picture: the subscription at the tier that includes the capabilities you need rather than the entry tier; any per-conversation, per-message, or per-resolution usage charges; the messaging channel's own fees, which on WhatsApp are billed per message by Meta and vary by message category and country; integration or onboarding fees; and the internal time to maintain the knowledge the agent is grounded in. Ask for a worked example at your volume, in writing.
Is an AI commerce agent worth it for a small store?
Not always, and a vendor who says otherwise is selling rather than advising. If your conversation volume is low enough that one person answers everything within a reasonable time, the agent's value is mostly in coverage outside working hours and in follow-up that would not otherwise happen. If your catalogue changes constantly and is not held in a system the agent can read, you will spend more time maintaining grounding than you save. Both are legitimate reasons to wait.
- AI agents
AI commerce platform architecture: from conversation to transaction
A reference architecture for AI commerce platforms: channels, identity, grounding, agent orchestration, tools, commerce, escalation, and evaluation — layer by layer.
- AI agents
AI agent guardrails: keeping a commerce agent on-brand
How to set guardrails that keep an AI commerce agent on-brand and accurate: grounding in your own content, a 'never improvise' list, and testing before launch.