You can now tell a computer to check which products are low on stock and take them off the front page, and it will go and do it. No developer, no integration, no ticket. It opens your admin in a browser and works through it the way a person would.
That became true in the first week of September, and it changes what an ecommerce platform is worth.
What the new models do
Two came out days apart, and neither headline was about the AI being smarter.
OpenAI released GPT-6 Astra on 3 September. What is new is that it can work a computer. It looks at the screen, moves the cursor, types into fields, clicks save, and keeps going for an hour without being told what to do next. OpenAI sells it on the dull half of everyone's job: filling in forms, updating records, working through long tasks with many steps.
Anthropic released Claude Fable 5.1 two days earlier, and that one was about money. The same kind of long, tool-heavy work now costs up to 45% less than it did.
Put together, one made it possible for software to be operated by something that is not a person, and the other made it cheap enough to run all day. That is the shift. Not a smarter chatbot. A worker.
For anyone running a store, the jobs it points at are obvious. Repricing 400 products for one market. Tagging a new drop into the right categories. Writing sixty product descriptions in a second language. Finding the products with one image where the rest have four. Building the sale page. None of it is hard. All of it takes hours, and the hours are why it waits.
Two ways to do the same job
There are two ways a machine can do that work. From the outside they look the same.
The first is the one described above. The AI uses your admin the way you do, reading the screen and clicking. It works on any platform, in any system, with no preparation at all, which is why it is getting all the attention.
The second is that the AI talks to your platform directly, through a list of jobs the platform allows: change this price, in this market, for this product. No screen involved. There is a common format for this now, called MCP, and the platforms have started to support it. Ask for the same thing and you get the same result.
The difference does not show on a good day. It shows on a bad one.
What happens when it goes wrong
The tests that measure this kind of work show a big jump on last year, and they also show that a fair share of long jobs still go wrong somewhere. That is on a test, not in your back office during campaign week.
So the question is not whether AI can reprice 400 products. It is what your store looks like on the day it repriced 180 of them and lost the thread.
If it worked through your screens, you get what a person would leave behind: a changed database and no explanation. A click tells the system what to do. It does not record what was meant, which market it was for, or which 180 of the 400 got there. Nobody signed anything. Nothing is easy to undo.
If it went through the platform directly, each change was a request the system understood: this product, this market, this price. It can be limited in advance to one market, shown to you before it goes live, listed afterwards and put back.
None of this assumes the AI is right. People are not right either. A merchandiser can tell you what they did, and a browser session cannot.
Fifteen years of selling the wrong thing
Ecommerce platforms have competed on their back office for as long as most of us have worked in this. How good the admin is. How fast a new hire learns it. How few clicks to a campaign. It is why replatforming hurts, and why "our team knows this system" wins arguments it should probably lose.
Every one of those advantages assumes a person at a screen.
There is something unfair buried in it too. The admin has always shown a fraction of what the platform can do. Somebody decided which parts a merchandiser should see and arranged them into a working day. For a person, that choosing is the service you are paying for. For a machine, it is a locked door in front of everything else.
So the question when you compare platforms stops being how nice the admin is, and becomes how much of the system can be reached without one.
What to ask a vendor now
The admin demo has stopped being the useful part of the meeting. Four questions do more:
- What can be done in this system without opening the admin, and is it everything the admin can do or a smaller set?
- How narrowly can access be limited? One market, one price list, one language, one category?
- Can a change be prepared and looked at before it goes live?
- Can you tell me afterwards exactly what changed, and put it back?
None of that is exotic. It is what you would want if you hired someone in September and gave them the keys on their first day.
Most of your work is not written down anywhere
The campaign process, the rule about which markets get the price change, the reason the third image is always the flat lay, the Monday routine someone worked out in 2023: almost none of it exists in writing. It lives as a habit in a person and comes out as clicks. It never had to be explained, because the person doing it already knew.
That is why handing work over is hard, and it was hard long before AI. Training a new merchandiser and instructing a machine turn out to be the same problem. If the only description of a job is watching someone do it, nobody can take it on. A company that has never solved this will not solve it by buying a subscription.
Writing it down is unglamorous and it is the work. What is this job called. What is it allowed to touch. Who looks at it before it goes live. How do we undo it.
What AI will not reach
Your ERP will not offer a tidy list of jobs. Neither will the warehouse system, the supplier portal, or the marketplace back office someone logs into twice a week. For all of that, an AI clicking through screens stays the only option, and that reach is a real advantage rather than a consolation.
It is also the risk. An AI driving your browser is logged in as you, with every market and every permission you have, and no way to say "this market only". Astra is the first model OpenAI rates at the top of its own scale for cybersecurity capability. That sits oddly with handing one your session and going to lunch.
Both will be true for years. The judgement worth making is which jobs deserve the careful path: the ones that touch price, stock, published content and customer data.
What the screens become
None of this empties the back office. It moves what the back office is for.
When typing things in is no longer the job, what is left is judgement. Seeing what the machine wants to do before it does it. Understanding what it touched. Catching the case a rule would have got wrong. Putting something back. That is a smaller surface than a full admin and a harder one to design, because everything on it has to help someone decide rather than help someone type.
We build storefronts on Centra and Shopify with design and engineering in one team, and we have written before about what AI can and cannot do in a build. The operations question is the same question one floor up, and it is worth asking while you are choosing rather than after.
For fifteen years the question was how good your back office is. The better question is how little of your work has to happen there.