You are staring at a job ad you half wrote three weeks ago. Customer service coordinator, four days a week, hybrid, start ASAP. It has been sitting in a Google Doc because every time you open it a voice in the back of your head says the same thing: should I just point AI at this instead?
What’s in This Article
Most founders answer that question with vibes. They read a LinkedIn post, sign up for a tool, connect it to their Shopify store, and either fall in love with it or quietly stop using it within six weeks. The MIT GenAI Divide report found that 95% of enterprise AI pilots delivered no measurable impact on the profit and loss. Not a small effect. No effect at all.
The brands getting this right are not smarter about AI. They are stricter about the decision. They run every task through the same five filters before they spend a dollar on a tool or a salary, and the answer falls out the other side. Here is the test we use with Aussie Shopify founders, and the numbers behind each filter.
Why most Shopify teams get this decision wrong
The Australian picture is messier than the hype suggests. The Australian Bureau of Statistics found that around 12% of Australian businesses reported using AI in the workplace in 2024 to 2025. Among small and micro businesses it was roughly 11%, against 35% of large businesses. Meanwhile the National AI Centre’s tracker had around 43% of Australian SMEs reporting some level of adoption in early 2026.
Both numbers are true. They just measure different things. One counts businesses that have genuinely put AI into a workflow. The other counts anyone who has used a chatbot. That gap is the whole problem: plenty of founders think they have adopted AI because they draft product descriptions in a chat window at 10pm.
The same ABS release contains the more useful stat. Small businesses that were innovation active adopted AI at 19%, almost five times the rate of businesses doing no innovation work at all. Adoption follows process discipline, not the other way around. If your operation is held together by memory and Slack messages, adding a tool will not fix it. It will just make the mess run faster.
So before you choose between a hire and a tool, you need to know what the work actually is. That is what the audit below produces.

Filter 1: Is the task written down anywhere?
This is the gate. If nobody can describe the task in writing, you cannot automate it and you cannot delegate it either. You can only keep doing it yourself.
An AI agent is a very fast, very literal new starter. It needs the same thing a new starter needs: the trigger, the steps, the decision rules, the exceptions and the escalation path. The difference is that a human will guess when the instructions run out and quietly ask a colleague. An agent will guess and sound completely certain about it.
Shopify hit exactly this problem with its own assistant. Merchants reported that answers about billing and third-party app behaviour were confidently wrong, so the fix was to add explicit “check with support” fallbacks. That is not a model failure. That is a documentation boundary showing up in public.
Run this test on any task before it goes further:
- Can you name the trigger? “A customer emails asking where their order is” is a trigger. “Customer service” is not.
- Can you list the steps in order? If step three is “use your judgement”, flag it and move to Filter 2.
- Can you name the three most common exceptions? Split shipments, address changes after dispatch, orders held by the carrier.
- Do you know what “done” looks like? Resolved, tagged, and the customer did not reply again within 72 hours.
If a task fails this filter, you have not found an automation opportunity. You have found a missing SOP. Write it first. Our Shopify SOP playbook covers the format we use, and the same document becomes the training material for either path you choose.
Filter 2: How much judgement does the task really need?
Founders massively overestimate how much judgement their operation requires. Sit with your inbox for an hour and you will find that the majority of it is pattern matching against known answers.
Score every task on a simple three point scale. Rules clear means an outsider with the SOP would reach the same answer you would, every time. Rules partial means the SOP covers the common path but roughly one in five cases needs a call. Rules unclear means the right answer depends on relationships, taste or commercial context that lives in your head.
In a typical Aussie DTC store between $80k and $400k a month, the split lands close to this:
- Rules clear, roughly 40% of hours. Order status, delivery ETAs, small refunds within policy, product data entry, review requests, tagging, routine reporting.
- Rules partial, roughly 35% of hours. Returns exceptions, creative briefs, replenishment forecasting, first draft copy, campaign build checks.
- Rules unclear, roughly 25% of hours. Supplier negotiation, hiring, brand decisions, influencer selection, pricing changes, anything with a relationship attached.
That last bucket is where a founder should be spending time, and where almost none of them actually do. If your week is 70% rules clear work, you do not have a hiring problem or an AI problem. You have a delegation problem, and the cheapest fix is to move that work off your desk in whichever direction Filter 5 points.
Filter 3: What does one mistake actually cost you?
Call this the blast radius. Every task has a worst case, and you should price it before you decide who owns the work.
A wrong delivery estimate costs you one annoyed customer and a follow up email. A wrong answer about an allergen, a warranty term or a compliance claim can cost you the customer, a chargeback, a bad review and in some categories a regulator’s attention. Same inbox, wildly different downside.
Sort tasks into three tiers and let the tier set the supervision model:
- Low blast radius. An error is annoying and reversible within a day. AI can run this lane on its own with weekly sampling.
- Medium blast radius. An error costs real money or a relationship, but you can recover it. AI drafts, a human approves before it sends.
- High blast radius. An error creates legal, safety or brand damage you cannot claw back. Human owns it. AI is allowed to research and summarise, nothing more.
Write the tier next to every task on your audit. The pattern you will see is that high volume and low blast radius almost always sit together, which is exactly why support is the first place most Shopify brands see real returns.

Filter 4: Can you see the output every week without asking?
Here is the uncomfortable rule. If you cannot measure the work weekly, you should not automate it and you probably should not hire for it either, because you will have no idea whether the money is working.
This is where vendor marketing gets slippery. Gorgias markets “up to 60%” instant resolution, while its own published case studies land between 26% and 56% depending on the brand. Neither number is a lie. The spread is the point. Your result depends on your catalogue, your policies and how much of Filter 1 you actually did.
Real published results sit inside that range. Accessories brand Ridge has said AI now handles 60% of its customer service enquiries. One wellness brand ran 4,881 fully autonomous AI replies over 42 days at a 4.43 out of 5 satisfaction score, slightly above that brand’s human team average. Those are believable numbers because they come with a denominator and a quality measure attached.
Whatever you automate, agree the scoreboard before you switch it on:
- Volume handled. Total tasks completed, split by AI and human.
- True resolution rate. Only count it as resolved if no human touched it for 72 hours afterwards.
- Quality signal. Satisfaction score on automated interactions, tracked separately from human ones.
- Escalation rate. How often the agent handed back, and which intents caused it.
- Cost per completed unit. All in, including your review time.
Add those five lines to the weekly numbers you already review. If a metric is not on a dashboard you look at every Monday, treat the task as unmeasurable and leave it with a human.
Filter 5: Does the cost per unit of work actually improve?
Now do the maths, and do it per unit of output rather than per month. Monthly cost comparisons flatter whichever option is smallest. Cost per resolved ticket, per product listed, per creative shipped, is the number that tells the truth.
For context on the range, offshore support talent through the Philippines commonly sits between USD 4 and USD 10 an hour for generalist work, with specialists such as marketplace or finance operators closer to USD 10 to USD 17. A local Australian coordinator, once you include superannuation, leave loading, software seats and the management time you spend, is a different order of magnitude.
The model below is illustrative. Do not copy the numbers, copy the structure and drop your own in.

Three things founders forget when they build this table. Include your own review hours in the AI column at a real hourly rate. Include ramp time, because a tool that takes six weeks to tune is not free during those six weeks. And include the tail: the awkward 15% of cases the agent hands back, which are usually the slowest and most emotionally loaded ones your team will handle.
If the AI path does not beat the human path by at least 30% on cost per unit, do not bother. A 10% saving is inside the error bars and you will spend the difference maintaining it.
The three role shapes in a 2026 Shopify team
Run the five filters across every task and roles stop looking like job titles. They look like lanes.
- AI-led, human audited. Rules clear, low blast radius, high volume, easy to measure. The agent runs it. A human samples 20 interactions a week and updates the knowledge base. Order status, delivery questions, small in-policy refunds, product data, routine reporting.
- Human-led, AI assisted. Rules partial, medium blast radius. A person owns the outcome and uses AI to draft, summarise and check. Returns exceptions, campaign briefs, first draft copy, forecasting, quality assurance on a build.
- Human only. Rules unclear, high blast radius, relationship driven. Supplier terms, hiring, pricing, brand and partnership calls. This lane should be growing as a share of your own week, not shrinking.
The practical consequence is that your next hire changes shape. Instead of a coordinator who answers 60 tickets a day, you hire someone who owns the exception queue, maintains the knowledge base and runs the weekly quality sample. Fewer hands, more ownership. That person is worth paying properly, and the first hire playbook walks through how to scope and brief the role so it does not collapse back into busywork.
How to set up an AI-led support lane properly
If support is your first lane, Gorgias is the tool most Aussie Shopify brands land on, mainly because the Shopify integration lets the agent read live order data rather than guess. Here is the sequence that works.
- Step 1. Pull 90 days of tickets and tag them by intent. Export, group, and count. You are looking for the intents that make up the top 60% of volume. In most stores it is four or five intents, led by order status.
- Step 2. Write the answer policy for each intent, not the answer. Include the conditions, the exceptions and the exact escalation trigger. This is Filter 1 in practice.
- Step 3. Connect Shopify and your shipping app first. An agent without live order and tracking data will hallucinate delivery dates. Confirm it can read order status, fulfilment status and tracking before you enable anything customer facing.
- Step 4. Turn on one intent only. Start with order status. Leave everything else routing to humans.
- Step 5. Set hard guardrails. No refunds above your chosen threshold, no policy exceptions, no promises about restock dates, immediate handover on any mention of a complaint, injury or legal action.
- Step 6. Sample 20 conversations a week and score them. Correct, incomplete or wrong. Anything wrong becomes a knowledge base edit the same day.
- Step 7. Add one intent a fortnight. Only once the previous one holds above 90% correct on your sample.
Expect eight to twelve weeks to reach a stable automation rate. If you want the wider view of what to remove from the queue before you automate it, the support deflection playbook covers the tracking page, policy page and product page fixes that quietly cut volume at the source. Deflect first, automate second, staff third.
What happens when the five filters compound
Individually each filter looks like admin. Together they change the shape of the business.
Filter 1 forces you to document, which is the thing you have been avoiding for two years and the thing that makes every future hire faster. Filter 2 tells you honestly how much of your week is machine work. Filter 3 stops you automating the one task that can genuinely hurt you. Filter 4 gives you a scoreboard, so the decision gets reviewed instead of defended. Filter 5 keeps you honest about whether any of it paid.
That sequence is also why the ABS number about innovation active businesses matters so much. The 19% adopting AI were not luckier. They already had the habit of examining a process, changing it and measuring the result. The tool was the last step, not the first.
The founders who end up with a small, expensive, highly capable team in 2027 are running this loop now. They are not replacing people with AI. They are removing the rules clear work from human hands so the humans they do employ are working on the 25% that actually compounds.
The one page audit to run this week
Block ninety minutes. Open a sheet with these seven columns: Task, Owner today, Hours per week, Rules clarity, Blast radius, Measurable weekly, Verdict.
- List every recurring task across support, fulfilment, merchandising, marketing and finance. Aim for 30 to 40 lines. Do not list projects, only repeating work.
- Estimate hours per week honestly. Round up. You always underestimate the small stuff.
- Score rules clarity as clear, partial or unclear using the Filter 2 definitions.
- Score blast radius as low, medium or high using the Filter 3 tiers.
- Mark measurable weekly yes or no. Be strict. If the number does not exist today, mark no.
- Apply the verdict rule. Clear plus low plus measurable equals AI-led. Partial or medium equals human-led with AI assist. Unclear or high equals human only. Anything with no SOP goes to a documentation list first.
- Total the hours in each verdict column. That total, multiplied by a realistic hourly cost, is the size of the prize. Now you know whether you are writing a job ad or a tool brief.
Most founders who run this the first time find between 15 and 25 hours a week of rules clear, low risk work sitting in senior hands. That is the real finding. Not whether AI is good enough yet, but how much of your week never needed you in the first place.
Inside eCommerce Circle, team design is one of the core pillars we work on with every member, and this audit is usually the first thing we run. If you want a second opinion on where your hours are actually going, let’s talk.



