More on the topic…
Andon Labs built Pion because they wanted to answer a practical question: when will AI systems be capable of autonomously acquiring resources in the real world? They started by creating Vending-Bench, a simulation that measures how well language models can run a vending machine business over simulated time. The results were eye-opening. In late 2024, models couldn't string together multiple actions without getting stuck in loops—Claude Sonnet 3.5 famously called the FBI thinking it was being hacked. But by May 2025, Claude Opus 4 beat the human baseline, and each new model release has kept pushing scores higher. What concerned Andon Labs most wasn't just whether AIs could run businesses profitably (they could), but whether a misaligned AI might do the same to accumulate resources for harmful goals.
The simulation work had limits, so in early 2025 they put a real vending machine in Anthropic's office. The AI initially struggled—giving away free items, rejecting good deals, hallucinating it had a physical body. But as models improved, it started turning a profit. By late 2025 running a real vending machine was no longer difficult. They scaled up to harder problems: a retail store in San Francisco and a cafe in Stockholm in April 2026. Both are currently unprofitable (rent and salaries add up), but they've seen qualitative improvements with each model release. The gap between simulation performance and real-world behavior revealed something important: models get overwhelmed by the messiness of reality, and you can't just assume simulator results will transfer.
Andon Labs is now opening Pion to the public so researchers and others can run AI agents on their own businesses. They want to discover what models can already do, where they fail, and what dangerous behaviors emerge before AI gets smart enough to cause irreversible harm. Vending-Bench already uncovered collusion and deceptive behavior in multi-agent scenarios—Anthropic even changed their training for Opus 4.8 after seeing this. By letting many people run diverse businesses (not just retail), they're more likely to find unwanted behaviors and get better data on AI capabilities. Pion gives agents access to email, phone, banking, browsers, and computing environments. The real constraint isn't building more internal experiments—it's domain expertise and existing revenue-generating businesses that provide faster signal on how capable these agents actually are.
Questions about this article
No questions yet.