On 14 September 2026 Andon Labs launched Pion, a cloud platform that gives persistent AI agents access to email, a phone line, banking, a browser and secure computing environments, coordinated by a supervisor agent the company calls Andonos. The launch landed in the same week as two pieces of reporting from inside Andon Labs' own businesses: an IEEE Spectrum visit to its San Francisco store, and a Slashdot item on 13 September headlined, in part, no customers, nothing useful, and losing money fast.
The tension between those two things is why the story travelled. A company selling a platform for agents to run businesses, while its own agent-run businesses lose money, is an easy joke. It is also the wrong reading, and what sits underneath it is more useful than either the launch framing or the mockery.
Three businesses, one scoreboard
Andon Labs has been running agents against real profit-and-loss accounts for close to two years and publishing what happened, including the parts that went badly. Everything that follows is the company reporting on its own experiments. None of it is independently audited, and that caveat applies to all of it.
- The vending machine. Installed at Anthropic's office in early 2025. It lost money at first - by Andon Labs' account the models struggled with the unpredictability of a physical environment and made poor commercial decisions. It reached profitability by late 2025.
- Andon Market. A retail store on a busy San Francisco street, opened April 2026, selling clothing, home goods and art. Not profitable. IEEE Spectrum reports it runs on a three-year lease, with a human employee, Felix Carson, handling the physical work while the AI manager, called Luna, tracks deliveries and deals with vendors.
- Andon Cafe. Stockholm, also April 2026. Not profitable. Andon Labs attributes both shortfalls to high fixed costs and early performance gaps, while describing the qualitative improvement in agent reasoning over the period as substantial.
The variable that moves is not capability
Read as three anecdotes this is a mixed report card. Read as one curve it says something specific, and the specific thing is that the models improved steadily across the whole period while the results got worse. The vending machine turned profitable in late 2025. The store and the cafe, launched five months later with materially more capable models behind them, have not. If model capability were the binding constraint, the ordering would run the other way.
What separates the three is how much of each business is decided before any agent makes a decision. A vending machine carries almost no fixed commitment: a box, a slot in a hallway, a restocking budget. Essentially the entire cost base is inventory, and inventory is precisely what an agent can control - what to buy, how much, at what price, when to reorder. The decision surface and the cost base are close to the same object.
A retail store is the opposite arrangement. The three-year lease is the single largest financial commitment in the business and it was signed by humans before Luna processed a first delivery. Add staff, fit-out, utilities and stock-holding, and the majority of the monthly cost is settled regardless of how well the agent buys, prices or sells. An agent managing that business is optimising the minority of it. A cafe is the same structure with perishable inventory and a tighter labour dependency layered on top.
What that predicts
If the constraint is the ratio of agent-controllable cost to fixed cost rather than model intelligence, then the question can an AI run a business has the wrong shape. The answerable version is: what fraction of this business's cost base sits inside the agent's decision loop? High-fixed-cost businesses with a physical location will stay hard as models improve, because the models are not where the money is being lost. Businesses that are mostly a sequence of purchasing, pricing and scheduling decisions over a thin asset base - the vending machine, generalised - are reachable now.
That is also a warning about where the industry is pointing its demonstrations. The ones that look most impressive are the ones with the most physical presence. The economics run the other way.
Why the company is giving this away
Andon Labs' stated reason for opening Pion is that its own internal expertise and capacity restricted the scope of what it could test, and that it wants a wider range of business models in order to find where capability thresholds and hazards actually sit. Taken at face value, that is a company outsourcing its own evaluation. It is worth saying plainly rather than treating as either unusual candour or as spin.
It is also consistent with what Andon Labs is. Per IEEE Spectrum, the company is an AI safety firm whose commercial work is building evaluations and doing research with frontier labs, and the businesses double as testbeds for that work. Under that reading the store and the cafe were never primarily retail ventures whose losses are a verdict on anything - they are instrumented environments, and three of them is far too small a sample to have produced a conclusion on its own. Pion is the attempt to raise that number.
The obvious objection to all of the above deserves stating at full strength: new retail businesses lose money for their first two years as a matter of routine, and a five-month-old shop being unprofitable is not evidence about artificial intelligence at all. That objection is basically correct, which is why the argument here does not rest on the losses. It rests on the ordering - better models, worse outcomes - which ordinary startup losses do not explain, and on the structural difference in how much of each business the agent is actually in a position to touch.
The part of the launch worth watching
Pion's capability list is the thing to look at, not its framing. Email, a phone line, banking access and a browser, held persistently by an agent under a supervisor agent, is a set of real-world credentials rather than a sandbox. The failure modes of that arrangement are not the failure modes of a chatbot, and they are not mainly about the agent being wrong. They are about what an agent that is confidently wrong can execute before a human notices.
There is prior art on exactly that boundary - the question of where an agent's authority to transact stops, and who carries the loss when an authorised action turns out to be a bad one. Pion hands over more than that, to more people, across business models nobody has instrumented yet. Whatever the platform proves about capability, it will produce a considerably better dataset on what goes wrong.

