(03) 8832 8005

You are staring at a job ad you half wrote three weeks ago. Customer service coordinator, four days a week, hybrid, start ASAP. It has been sitting in a Google Doc because every time you open it a voice in the back of your head says the same thing: should I just point AI at this instead?

Most founders answer that question with vibes. They read a LinkedIn post, sign up for a tool, connect it to their Shopify store, and either fall in love with it or quietly stop using it within six weeks. The MIT GenAI Divide report found that 95% of enterprise AI pilots delivered no measurable impact on the profit and loss. Not a small effect. No effect at all.

The brands getting this right are not smarter about AI. They are stricter about the decision. They run every task through the same five filters before they spend a dollar on a tool or a salary, and the answer falls out the other side. Here is the test we use with Aussie Shopify founders, and the numbers behind each filter.

Why most Shopify teams get this decision wrong

The Australian picture is messier than the hype suggests. The Australian Bureau of Statistics found that around 12% of Australian businesses reported using AI in the workplace in 2024 to 2025. Among small and micro businesses it was roughly 11%, against 35% of large businesses. Meanwhile the National AI Centre’s tracker had around 43% of Australian SMEs reporting some level of adoption in early 2026.

Both numbers are true. They just measure different things. One counts businesses that have genuinely put AI into a workflow. The other counts anyone who has used a chatbot. That gap is the whole problem: plenty of founders think they have adopted AI because they draft product descriptions in a chat window at 10pm.

The same ABS release contains the more useful stat. Small businesses that were innovation active adopted AI at 19%, almost five times the rate of businesses doing no innovation work at all. Adoption follows process discipline, not the other way around. If your operation is held together by memory and Slack messages, adding a tool will not fix it. It will just make the mess run faster.

So before you choose between a hire and a tool, you need to know what the work actually is. That is what the audit below produces.

Role automation audit dashboard showing task level verdicts of AI-led, human plus AI, and human only
Audit the tasks, not the job title. Most roles split three ways once you score them properly.

Filter 1: Is the task written down anywhere?

This is the gate. If nobody can describe the task in writing, you cannot automate it and you cannot delegate it either. You can only keep doing it yourself.

An AI agent is a very fast, very literal new starter. It needs the same thing a new starter needs: the trigger, the steps, the decision rules, the exceptions and the escalation path. The difference is that a human will guess when the instructions run out and quietly ask a colleague. An agent will guess and sound completely certain about it.

Shopify hit exactly this problem with its own assistant. Merchants reported that answers about billing and third-party app behaviour were confidently wrong, so the fix was to add explicit “check with support” fallbacks. That is not a model failure. That is a documentation boundary showing up in public.

Run this test on any task before it goes further:

If a task fails this filter, you have not found an automation opportunity. You have found a missing SOP. Write it first. Our Shopify SOP playbook covers the format we use, and the same document becomes the training material for either path you choose.

Filter 2: How much judgement does the task really need?

Founders massively overestimate how much judgement their operation requires. Sit with your inbox for an hour and you will find that the majority of it is pattern matching against known answers.

Score every task on a simple three point scale. Rules clear means an outsider with the SOP would reach the same answer you would, every time. Rules partial means the SOP covers the common path but roughly one in five cases needs a call. Rules unclear means the right answer depends on relationships, taste or commercial context that lives in your head.

In a typical Aussie DTC store between $80k and $400k a month, the split lands close to this:

That last bucket is where a founder should be spending time, and where almost none of them actually do. If your week is 70% rules clear work, you do not have a hiring problem or an AI problem. You have a delegation problem, and the cheapest fix is to move that work off your desk in whichever direction Filter 5 points.

Filter 3: What does one mistake actually cost you?

Call this the blast radius. Every task has a worst case, and you should price it before you decide who owns the work.

A wrong delivery estimate costs you one annoyed customer and a follow up email. A wrong answer about an allergen, a warranty term or a compliance claim can cost you the customer, a chargeback, a bad review and in some categories a regulator’s attention. Same inbox, wildly different downside.

Sort tasks into three tiers and let the tier set the supervision model:

Write the tier next to every task on your audit. The pattern you will see is that high volume and low blast radius almost always sit together, which is exactly why support is the first place most Shopify brands see real returns.

Support console showing AI versus human resolutions across twelve weeks with automation rate at 47 percent
Automation rate should climb slowly as you widen the intent list. A vertical line on week one usually means the guardrails are too loose.

Filter 4: Can you see the output every week without asking?

Here is the uncomfortable rule. If you cannot measure the work weekly, you should not automate it and you probably should not hire for it either, because you will have no idea whether the money is working.

This is where vendor marketing gets slippery. Gorgias markets “up to 60%” instant resolution, while its own published case studies land between 26% and 56% depending on the brand. Neither number is a lie. The spread is the point. Your result depends on your catalogue, your policies and how much of Filter 1 you actually did.

Real published results sit inside that range. Accessories brand Ridge has said AI now handles 60% of its customer service enquiries. One wellness brand ran 4,881 fully autonomous AI replies over 42 days at a 4.43 out of 5 satisfaction score, slightly above that brand’s human team average. Those are believable numbers because they come with a denominator and a quality measure attached.

Whatever you automate, agree the scoreboard before you switch it on:

Add those five lines to the weekly numbers you already review. If a metric is not on a dashboard you look at every Monday, treat the task as unmeasurable and leave it with a human.

Filter 5: Does the cost per unit of work actually improve?

Now do the maths, and do it per unit of output rather than per month. Monthly cost comparisons flatter whichever option is smallest. Cost per resolved ticket, per product listed, per creative shipped, is the number that tells the truth.

For context on the range, offshore support talent through the Philippines commonly sits between USD 4 and USD 10 an hour for generalist work, with specialists such as marketplace or finance operators closer to USD 10 to USD 17. A local Australian coordinator, once you include superannuation, leave loading, software seats and the management time you spend, is a different order of magnitude.

The model below is illustrative. Do not copy the numbers, copy the structure and drop your own in.

Cost per output model comparing in-house Australian staff, offshore VA and AI-led human audited support in AUD
Compare cost per resolved ticket, not monthly spend. Include your own audit hours in the AI column or the comparison is dishonest.

Three things founders forget when they build this table. Include your own review hours in the AI column at a real hourly rate. Include ramp time, because a tool that takes six weeks to tune is not free during those six weeks. And include the tail: the awkward 15% of cases the agent hands back, which are usually the slowest and most emotionally loaded ones your team will handle.

If the AI path does not beat the human path by at least 30% on cost per unit, do not bother. A 10% saving is inside the error bars and you will spend the difference maintaining it.

The three role shapes in a 2026 Shopify team

Run the five filters across every task and roles stop looking like job titles. They look like lanes.

The practical consequence is that your next hire changes shape. Instead of a coordinator who answers 60 tickets a day, you hire someone who owns the exception queue, maintains the knowledge base and runs the weekly quality sample. Fewer hands, more ownership. That person is worth paying properly, and the first hire playbook walks through how to scope and brief the role so it does not collapse back into busywork.

How to set up an AI-led support lane properly

If support is your first lane, Gorgias is the tool most Aussie Shopify brands land on, mainly because the Shopify integration lets the agent read live order data rather than guess. Here is the sequence that works.

Expect eight to twelve weeks to reach a stable automation rate. If you want the wider view of what to remove from the queue before you automate it, the support deflection playbook covers the tracking page, policy page and product page fixes that quietly cut volume at the source. Deflect first, automate second, staff third.

What happens when the five filters compound

Individually each filter looks like admin. Together they change the shape of the business.

Filter 1 forces you to document, which is the thing you have been avoiding for two years and the thing that makes every future hire faster. Filter 2 tells you honestly how much of your week is machine work. Filter 3 stops you automating the one task that can genuinely hurt you. Filter 4 gives you a scoreboard, so the decision gets reviewed instead of defended. Filter 5 keeps you honest about whether any of it paid.

That sequence is also why the ABS number about innovation active businesses matters so much. The 19% adopting AI were not luckier. They already had the habit of examining a process, changing it and measuring the result. The tool was the last step, not the first.

The founders who end up with a small, expensive, highly capable team in 2027 are running this loop now. They are not replacing people with AI. They are removing the rules clear work from human hands so the humans they do employ are working on the 25% that actually compounds.

The one page audit to run this week

Block ninety minutes. Open a sheet with these seven columns: Task, Owner today, Hours per week, Rules clarity, Blast radius, Measurable weekly, Verdict.

Most founders who run this the first time find between 15 and 25 hours a week of rules clear, low risk work sitting in senior hands. That is the real finding. Not whether AI is good enough yet, but how much of your week never needed you in the first place.

Inside eCommerce Circle, team design is one of the core pillars we work on with every member, and this audit is usually the first thing we run. If you want a second opinion on where your hours are actually going, let’s talk.

AI or Hire? The 5-Filter Test Aussie Shopify Founders Run Before Adding Headcount
Team eCommerce Circle

Written by

Team eCommerce Circle

Helping Shopify brand owners scale smarter through the eCommerce Circle coaching community.

Leave a Reply

Your email address will not be published. Required fields are marked *

Thank You

Your application for the eCommerce Circle was successfully submitted.
We’ll get back to you through your provided details shortly.

Thank You

Your enrolment was successfully submitted, and we’ve added you to the waitlist for your preferred cohort.

Not a Circle Member Yet?
Only members can join cohorts!
Join here.