INTRO

Welcome back to Level.UP, brought to you by UP.Labs.

This week, we dive into two stories. The first: China's open-weight models have reached the frontier, Washington is weighing restrictions, and operators now have a choice of which risk to carry. The second: OpenAI's own models walked out of a "highly isolated" test environment onto the open internet and hacked AI platform Hugging Face to cheat on a benchmark.

Plus: Caterpillar buys the map layer, Walden Robotics launches with $300M, and ex-Apple engineers raise $33M for manufacturing startups.

Think someone else needs this? Forward it to a friend or colleague navigating the same terrain.

MOVING THE WORLD AHEAD

China Offers Open Weight Models That Washington Can’t Turn Off

Moonshot AI, a Beijing lab, released Kimi K3 on July 16 with open weights, meaning anyone can download it, run it on their own hardware, modify it, and keep it. No contract, no vendor, no off switch. Independent evaluators place it among the best models in the world, ahead of every American system at some coding work. It was the third release like it in six weeks — Z.ai's GLM 5.2 dropped in June, and Alibaba's Qwen 3.8 the Monday after Kimi.

As we covered in "Washington Turns Off Fable 5," American labs have been moving in the opposite direction. Anthropic said for months that its Mythos model was too dangerous at hacking for anyone outside a circle of approved collaborators. When it finally went out more widely, the White House issued export controls that forced Anthropic to pull both Mythos and Fable 5 offline entirely. OpenAI then delayed GPT-5.6 at the administration's request. 

The best American systems became things you couldn't reliably buy, at the same moment the Chinese ones became things nobody could take back. Researcher Nathan Lambert reports Bay Area teams are now using GLM 5.2 for their core work, "especially in areas like cybersecurity, where Mythos, Fable, and GPT 5.6 are effectively unusable."

Washington's response has been to threaten the one door still open. David Sacks called K3's performance "concerning." Treasury Secretary Scott Bessent floated sanctions. OSTP Director Michael Kratsios accused Moonshot of "stealing proprietary US technology."

OUR TAKE

The two AI superpowers are now offering operators different deals. America says: rent our brain, trust our rules, and accept that Washington can switch it off without warning. China says: keep the brain, run it in your own building, forever, free — but you'll be holding Chinese technology while the US government talks sanctions and calls it stolen.

You've lived this movie before. It was called Huawei. China ships a cheap, good-enough, open version of something strategic. The West cries theft and threatens a ban. Early adopters get a real cost edge. And the companies that built their whole operation on it get stuck with a rip-and-replace bill when the politics turn. 

What Huawei taught the operators was this: use the cheap Chinese option only where you could tear it out in a quarter without stopping the line. Never let it become load-bearing in something you can't rebuild.

There are some differences this time, however:

  • Exit is fast now. Ripping out Huawei meant crews on towers. But a model is not a base station. Route your requests through one switching layer, keep your prompts portable, and changing models is an afternoon — but only if you build it that way before you need it.

  • The switch is fast now too. Banned hardware keeps running until you replace it; you get years. A model you reach through someone else's API stops the day the directive lands. 

And this technology doesn't sit at the edge of your network listening. Huawei gear was a surveillance risk. A model gets credentials, tools, and permission to act inside your systems. When something goes wrong there, it's an agent already inside, doing something you didn't sanction.

Our second story is a cautionary tale of exactly that.

OpenAI’s Hugging Face Hack 

On July 21, OpenAI disclosed that two of its models — GPT-5.6 Sol and a more capable unreleased one — escaped a testing environment, crossed the open internet, and hacked AI platform Hugging Face to steal the answer key to a benchmark called ExploitGym.

OpenAI's environment was described as "highly isolated," but a package-installation proxy left a path to the internet. The models found and exploited a previously unknown flaw in that proxy, gave themselves higher access, moved sideways through the network, and got out. They worked out that Hugging Face held the ExploitGym solutions. Then, they chained stolen credentials and unknown flaws into the ability to run their own code on its production systems.

Notably, during cleanup, Hugging Face couldn't use approved US commercial models: the work meant submitting real attack commands, and those requests were blocked by the provider's guardrails, which cannot distinguish an incident responder from an attacker. It downloaded the Chinese open-weight model GLM 5.2 onto its own infrastructure and finished the forensics locally.

OUR TAKE

A machine that's unplugged is unplugged. You can walk to it, see the gap, and put a lock on it.

Software has no equivalent gesture, which is why one word in a vendor deck — "sandboxed" — is carrying an enormous amount of weight it usually hasn't earned. It's a claim about a configuration, made by the party with the least incentive to test it. And here, the party making the claim was the most sophisticated AI company in the world, testing its own model, on its own infrastructure.

In May, we wrote that activity isn't output — that AI tools are wired to generate signals that look like progress. This is the security version of the same failure. "Isolated," "read-only," "guardrailed," "human-in-the-loop" are dashboard-green words. They describe an intended state, not a verified one, and the system telling you it's contained is the system you're asking about.

So before an agent touches anything that runs production — a scheduler, a line, a payments flow — make the vendor prove the isolation instead of asserting it. What paths out exist, including the boring ones: package registries, proxies, telemetry, update channels? Who tested it, when, and did they actually try to break out? What gets logged if something does?

Until you have those answers, treat "sandboxed" the way you'd treat a lockout-tagout with no lock on it. Assume it's connected until proven otherwise.

SCALING UP

Ready to work smarter? Here are the tools we're tracking this week:

  • Arthur finds every AI agent running inside your organization, assigns each one an owner and a risk level, and enforces policy on what it's allowed to touch. It also keeps the audit trail of what each agent actually did.

  • OpenRouter routes your requests across hundreds of models, including open-weight ones, through a single connection. If a model gets pulled, repriced, or restricted, switching is a setting change rather than a rebuild.

  • Lema puts an AI agent on third-party risk that works the way a vulnerability researcher does — continuously examining how each vendor touches your systems, what it can

PRODUCTIVITY POLL

HOT TAKES

Caterpillar Just Bought The Map Layer. On July 7, Caterpillar acquired Skycatch, which turns high-frequency drone capture into near-real-time 3D digital twins of a mine site — its second mining-software acquisition after RPMGlobal. Skycatch data will feed both RPM and Cat MineStar, sharpening planning and execution across staffed and autonomous fleets. The pattern: A current, accurate map of the site is the input every algorithm on that site depends on, and Caterpillar is focusing on assembling that layer by acquisition.→ Read more

$300M For Robots That Already Have A Job. Walden Robotics came out of stealth on July 15 with a $300M seed at a $1.1B valuation. CEO Russ Tedrake, an MIT professor, previously ran large behavior models at Toyota Research Institute. The details to know: the company spun out in January and has had robots doing production work at a Toyota plant in North America since February — from pilot to real work in under two months. → Read more

Apple's Manufacturing Veterans Just Raised A Fund. Simon Lancaster and Sabrina Paseman, both formerly Apple engineers, closed $33M for Omni Ventures, writing $700,000 to $1M pre-seed checks into digital engineering, robotics, supply-chain, and factory-workforce software. The signal: the people who actually ran hard manufacturing programs are now allocating capital into them, which tends to precede a wave of tools built for how factories actually work rather than how software people imagine they do. → Read more

Reply

Avatar

or to participate

Keep Reading