This week, I read about a study at the University of Zurich in which scientists recruited 1,000 volunteers to bury 2,000 pairs of cotton underwear across Switzerland.
I promise this is going somewhere useful.
The researchers wanted to know how healthy the soil was in different places. Healthy soil is full of tiny living things that eat and break down cotton.
The microbes doing the work were underground, so researchers judged their activity by what they left behind. They dug up the underwear after two months. The more it had rotted, the more active the soil.
Reading about the study, I kept thinking about the newly released GPT-6 Astra. It can do more work on its own but reveals less about how it reaches a result. We may need to assess AI the way those researchers assessed the soil: by looking at what it leaves behind.
That also made me wonder: how do you trust an AI you can no longer watch think?
Astra, the more autonomous model that explains itself less
OpenAI released its newest model, GPT-6 Astra, on September 3. OpenAI president Greg Brockman said it likely represents artificial general intelligence: AI that's broadly capable across different tasks, not just good at whatever it was trained hardest on.
The benchmark results are less dramatic. On a benchmark built to test reasoning in unfamiliar situations, OpenAI announced a 99.9% score using its own specialized setup. Run the standard way, independently, that score drops to 62.7%. On that test, Astra improved a lot over the model it replaced (Sol at 7.8%).
Its autonomy changed too. Astra can browse the web, run programs, and complete long tasks without a person checking each step. But OpenAI's own safety testing found its reasoning is now "harder to monitor," and just knowing it was being watched was enough to make the model explain itself less.
This is the same model OpenAI delayed last month over concerns it had gotten too good at hacking on its own. It shipped anyway, then aced a test built to find and exploit software vulnerabilities. The intelligence didn't jump. The autonomy did. The transparency dropped.
It's also gotten much cheaper to use a capable model. A similar task that cost OpenAI about $4,560 to solve well in December 2024 can now be solved even better for under a dollar.
That price drop makes it easier for employees to use autonomous AI without requesting a new budget or going through procurement. If access controls are weak, they may also connect it to company data and tools before anyone reviews the risks.
Why this reminds me of the buried-underwear problem
I did not expect to spend part of this week thinking hard about decomposing underpants, but here we are.
You cannot watch a chain-of-thought that is hidden from view.
So you check the same way those Swiss researchers did. Skip the process. Go straight to the result.
Most companies already verify at least some AI output. The change is that the model's explanation is becoming a less reliable source of assurance.
You can still monitor an AI system through logs, access controls, behavioral tests, alerts, and output reviews.
As visible reasoning becomes less useful, those checks need to carry more weight. Occasional spot checks are not enough for systems that can act without approval.
Different ways to verify AI’s results

The four practical checks I swear by:
- Make people commit to their own answer before they see the AI's. Anchoring to the AI's answer first can erode judgment over time.
- Run high-stakes output through a second, unrelated AI model and ask it to find flaws. Different models often fail in different ways. If they disagree, send the output to a person for review.
- Occasionally plant a wrong answer in your review queue. This tests the reviewers, not the AI. If nobody catches it, strengthen the review process even if the AI also needs work.
- Track "confidently wrong" as its own category, separate from "wrong." Record whether the AI expressed uncertainty. Over time, this will show which kinds of tasks produce confident errors, so you can review similar outputs more closely in the future. You still need to verify each output to know whether it was wrong.
I've started doing something small with my team I never bothered with before: comparing the AI's draft against the finished, cleaned-up version, even on tasks that should be simple.
If the final version needs factual corrections, missing context, or a major rewrite, we review the task instructions, source material, examples, and approval steps.
It takes about fifteen minutes a week and tells us what to change about our process and inputs before the next draft.
The jobs AI is actually creating
Verification takes real hours every week. Companies must decide where those hours will come from.
In the US, AI has created roughly a million jobs while cutting about two hundred thousand, according to The Economist.
Most of that growth is in building and running AI, plus construction for the data centers and power grid behind it. That creates work for electricians, HVAC technicians, and grid engineers.
I am not hiring data-center electricians, and I doubt you are either. The economy-wide numbers are real, but they do not tell a single company what to do with its people.
For an enterprise LMS company like ours, it comes down to two questions: do you automate, and what do you do with the labor that frees up?
My answer to the first one is easy. Yes.
Start with tasks that are repetitive, reversible, and measurable, where errors are easy to detect and correct. Some of those tasks may sit with junior employees, but seniority is not the test.
What to do with the time AI frees up: a lesson from Toyota
Use the time saved through automation to improve how the work gets done. Lean and Six Sigma have applied that discipline for decades. Remove work that does not serve the customer, then use the saved time to find and fix the next problem.
Some companies will keep all the savings from automation. Others will use part of the saved time to prevent recurring work.
Take customer support. Automating routine tickets doesn't have to mean answering more tickets faster with the same team. It can mean giving someone time to find out why customers keep filing the same three tickets, then fixing the product so those tickets stop.
I know which of those two I want my company to be.
Toyota ran a real version of this. During the 2008–2009 recession, as US auto sales fell to lows not seen since the early 1980s, Toyota idled plants in San Antonio and Princeton, Indiana rather than laying off roughly 4,500 workers.
It kept them on full pay and put their time into training and kaizen projects. At the Princeton plant, one project found that installing door padding required 23 kilograms of push force, damaging material and straining workers. The fix brought it down to 6.8 kilograms, cutting defects and saving about $6,000 a month in scrap and rework.
Toyota cited employment security as a reason for avoiding layoffs. Employees have more reason to suggest improvements when they do not fear that those improvements will cost them their jobs.
Companies can apply the same principle to AI. Use saved time to check outputs, trace recurring errors, and improve the instructions, data, permissions, and review steps around each system. Employees who manage AI-supported processes will spend more of their time on that work.
What to do this week
Astra still leads plenty of leaderboards. Oversight has not become impossible, but the model's own explanation provides less assurance.
At the same time, the model can take more autonomous action. Logs, permissions, behavioral tests, alerts, and output audits now matter more than ever.
This week, I challenge you to pick one production process in which AI drafts, decides, or acts, but nobody has audited recent outputs.
Ask who samples those outputs, what they check, and how often.
If no one owns the review, assign an owner and set a schedule.
What's something AI got wrong for you this week? And how did you work to stop it from happening again?

.avif)
.avif)
.avif)