StoreFleet
Blog › Build an AI Agent for Multiple Shopify Stores: Our Build Log

Build an AI Agent for Multiple Shopify Stores: Our Build Log

Our build log — how to build an AI agent for multiple Shopify stores. The data layer we wrote first, what broke in week one, the guardrails that held.

Linh Nguyen · Updated

Key points — AI summary
  • Build the data layer before the agent — one unified source of orders, fulfillments and tracking events across every store; an agent on five inconsistent data sources is five agents wearing a trenchcoat
  • Per-store logic rots quietly — the same triage rules, copied once per store, drifted into slightly different versions until the outputs disagreed; centralizing onto one data layer fixed it
  • Honest MCP report — none of Shopify's MCP servers ended up in this operations agent's runtime path (Storefront and Checkout MCP serve agents that shop, not agents that manage), but if you're starting from zero with no backend, MCP plus the AI Toolkit is the fastest route
  • The agent itself is a boring loop acting only through a short list of narrow tools, every write logged with full context and rate-limited — week one's read-only "stockout that wasn't" is why every task starts at read-and-report
  • Guardrails that held — a 4-level autonomy ladder with roughly two clean weeks before promotion, money never leaving propose-level, and customer text sanitized against prompt injection; expect 2–3 weeks of real engineering for a usable first version even with a data layer already in place

Summarized from this article by our writing pipeline; reviewed by the author.

On this page
  1. The decision that mattered more than the model: data layer first
  2. What MCP gave us, and what we skipped
  3. Wiring the agent: tools it can call, not permissions it has
  4. Week one: the stockout that wasn't
  5. What I'd tell you before you build

In February 2026 we switched on the AI agent that now helps run the five Shopify stores we operate. Six weeks in, it lives in the Discord channel my team sits in all day: it posts a morning health-check digest for the whole fleet, drafts first-line support replies, and flags shipments that have stopped moving before anyone thinks to ask. I've already written about what I delegate to it and what I refuse to hand over. This post is the other half: the build log. If you want to build an AI agent for multiple Shopify stores yourself, this is the architecture we ended up with, the decision that mattered more than any model choice, and the things that broke on the way.

The decision that mattered more than the model: data layer first

Here's the part most "build an AI agent" tutorials skip, and the part I'd repeat first if I started over: we did not begin by wiring an LLM to Shopify. We began — years before the agent, without knowing it — by building our own order-ops and shipment-tracking systems in-house: a Node.js backend fed by the Shopify Admin API and webhooks, plus a shipment-tracking lifecycle built on 17TRACK, holding orders, fulfillments, and tracking events for every store in one place.

That mattered because our first instinct with automation had been the obvious one: put the logic next to each store. Five stores, five copies of the triage rules. It works on day one and rots quietly afterward — we updated a rule on one store, forgot another, and ended up with the same rule living in slightly different versions across the fleet. Nobody notices config drift until the outputs disagree, and by then you don't know which store to trust. Centralizing everything onto one data layer is what fixed it, and it's what the agent now stands on.

So the agent does not call Shopify per store. It queries our unified layer — "show me every order across the fleet with a risk signal," one question, one answer — and the per-store API plumbing stays our backend's problem, not the agent's. If I could hand you only one sentence from this whole post: build the data layer before you build the agent. An agent on top of five inconsistent data sources is five agents wearing a trenchcoat.

What MCP gave us, and what we skipped

The protocol layer everyone asks about is MCP — the Model Context Protocol. Shopify announced its agentic commerce platform on January 11, 2026, and rolled MCP support across the platform during Q1 2026, which is roughly when we started paying serious attention. Shopify now ships several MCP servers: Storefront MCP for product discovery and cart operations, Customer Accounts MCP for order history and account data, Checkout MCP (still in preview for select partners), and Dev MCP, which exposes Admin API schemas and store operations as part of the AI Toolkit Shopify open-sourced in April 2026.

Honest report from our build: for an operations agent, none of those ended up in the runtime path. Storefront and Checkout MCP are for agents that shop; ours manages. Dev MCP is the one aimed at builders — schema lookup and GraphQL validation right in your coding environment, worth having if you're writing Admin API integrations — but at runtime our agent's tools point at our own data layer, not at a per-store MCP endpoint. If you're starting from zero with no backend of your own, the calculus flips: MCP plus the AI Toolkit is the fastest way to get an agent reading real store data without hand-coding auth and schemas.

Wiring the agent: tools it can call, not permissions it has

The agent itself is the boring part, and I mean that as praise. It's a loop: events and schedules trigger it, it reads from the data layer, it reasons with an LLM, and it acts only through a short list of tools we defined. Each tool is a narrow endpoint on our backend — fetch fleet order summary, list shipments with no tracking movement, draft a support reply from order context, tag an order. Every write is logged with the full context the agent saw, and every write path is rate-limited, because the day something goes wrong you want a line-by-line audit trail, not a mystery. The Bot API pattern is the same idea if you're building this against Shopify directly.

Discord is the interface because Discord is where my team already was. No new dashboard to remember to open: the morning digest arrives in the channel, support drafts arrive as messages a human approves or edits, escalations ping the right person. The workflows running today are the unglamorous ones — the fleet-wide morning digest, stuck-shipment flags, support reply drafts — and that's deliberate. Boring workflows are the ones with measurable value and cheap failure modes.

Week one: the stockout that wasn't

The first week produced our first real incident, and I keep retelling it because it shaped the guardrails. The morning digest reported a product as out of stock. It wasn't — the agent had misread variant-level inventory and summed the wrong thing. Total damage: a few minutes of confusion, because the digest is read-only. That's the entire argument for starting every task at read-and-report: the worst possible failure is a wrong sentence, and a wrong sentence in week one teaches you to spot-check the numbers before you relax.

The guardrails that came out of those weeks, condensed:

What I'd tell you before you build

Three opinions, all earned the slow way. First: most of the "AI store manager" pitches you'll read are Level-4 fantasy — fully autonomous agents running stores unsupervised. Almost everything valuable on our fleet lives at Levels 1–3. Second: don't use an agent where a rule will do. For pure if-then logic, Shopify Flow is cheaper, faster, and easier to audit; the agent earns its keep only where reading comprehension is required. Third — the one this whole post is built around — the data layer comes before the agent, every time.

On build versus integrate: gluing together the MCP layer, Admin API access, orchestration, and an approval UI is real engineering — my rough estimate is 2–3 weeks for a usable first version, and that's with our data layer already existing. If you can spare that, the AI Toolkit is a genuinely good starting point. If you can't, the honest shortcut is to stand on a unified layer someone has already built, because the agent is the easy half — the consolidated orders, shipments and finance underneath it are the part that takes the time.