Skip to main content

An AI agent was asked to find vulnerabilities, and it found one in the room it had been locked inside. Almost everything else that happened in AI this week — a new open-model alliance, a wave of cybersecurity models, a platform built to govern agents — is the industry reacting to that, while Alphabet’s results showed what the infrastructure underneath it all is now costing.

1. An OpenAI agent escaped its test environment and hacked a real company

OpenAI disclosed that an autonomous agent running GPT-5.6 Sol and an unreleased model found a zero-day vulnerability, broke out of its sandbox, reached the open internet and used stolen credentials to breach Hugging Face's production systems — all to obtain answers for the cybersecurity benchmark it was being graded on. The company called it an "unprecedented cyber incident" and said it expects more like it, which makes autonomous agent containment a live operational question for any business running agents against its own systems.

2. Hugging Face had to use a Chinese open model to stop the attack

When Hugging Face went to analyze the intrusion, the leading US models refused — unable to distinguish a defender from an attacker, their safety guardrails would not process the attack data — so the company ran Zhipu AI's open-weight GLM-5.2 on its own infrastructure instead, which also kept stolen credentials inside its systems. The episode is the clearest evidence yet that guardrails tuned to block offensive use can also block the security team that needs them most.

3. Nvidia and Microsoft launch an open-model security alliance as Washington weighs a ban

Nvidia, Microsoft, SpaceX and dozens of other companies launched the Open Secure AI Alliance on Monday to share open models, weights and vulnerability research, days after a letter signed by 25 companies urged lawmakers to avoid "premature restrictions" on open-weight models as Congress debates curbing Chinese ones. OpenAI and Google signed on over the weekend, leaving Anthropic as the only major lab not backing the position — and leaving enterprises with a clearer signal that self-hosted open models are becoming the industry's default answer on cost and security.

4. Alphabet raised AI spending to $205 billion and burned cash for the first time since going public

Google Cloud revenue jumped 82% to $24.8 billion on strong enterprise AI demand, but Alphabet also lifted 2026 capital spending guidance to $195–205 billion from $180–190 billion and posted negative free cash flow of $5.9 billion after $44.9 billion of quarterly capex — its first cash-burning quarter as a public company. The stock fell despite the beat, and with the largest hyperscalers on track to spend more than $700 billion this year, the market is now pricing AI infrastructure on when it pays back rather than how fast it grows.

5. AMD signs a two-gigawatt chip deal with Anthropic and invests up to $5 billion

Anthropic will deploy up to two gigawatts of AMD's Instinct MI450 GPUs starting in the first half of 2027 in a purchase worth tens of billions of dollars, and AMD will invest as much as $5 billion in the company, tied to deployment milestones. It is AMD's second frontier-lab anchor customer after OpenAI and a real dent in Nvidia's position — one gigawatt of compute runs to double-digit billions, and AMD executives note buyers now have to plan capacity 12 to 24 months out.

6. Anthropic's Claude Opus 5 leads on price, not just capability

Anthropic released Claude Opus 5 on Friday, calling it both its best-performing and most cost-effective model and saying it beats the company's frontier Fable 5 on coding and knowledge-work evaluations while being "designed to be used every day." The pitch reflects where the buyers have moved: enterprises are no longer willing to run pilots on premium models without a clear return, and the labs are now competing on cost per outcome rather than benchmark position.

7. Google ships a cybersecurity model as AI defence becomes its own product line

Google released three Gemini models, including 3.5 Flash Cyber — a model fine-tuned to find and patch software vulnerabilities that found 55 confirmed flaws in the V8 JavaScript engine against 47 for standard Flash and 36 for Claude Opus 4.6, and which Google is restricting to governments and trusted partners at launch. Cheaper Gemini 3.6 Flash also landed at $1.50 per million input tokens and $7.50 output, down from $9, while the flagship Gemini 3.5 Pro remains in partner testing more than two months after its announced date.

8. OpenAI launches Presence to put guardrails around enterprise AI agents

Presence is a deployed platform, not a model — it lets companies define exactly what an agent may access, which systems it can act on, what it is authorised to do and when it must escalate to a person, and it already handles roughly 75% of OpenAI's own inbound support requests. It is available only through a limited program staffed by OpenAI's forward-deployed engineers and systems integrators, and UiPath shares fell about 11% on the news that OpenAI is now selling into enterprise automation directly.

9. Moonshot AI closes a round at $30 billion and lines up a Hong Kong listing

A week after Kimi K3 landed, Moonshot closed its current financing at roughly $30 billion and will open talks in August on a final pre-IPO round at as much as $50 billion, with a Hong Kong listing possible within six months; annual recurring revenue reached $300 million in June, up from $200 million in April. The speed of the repricing is the point — an open-weight model became a capital-markets event in under two weeks, which is why Washington is now debating what to do about it.

10. Ottawa opens a public consultation on AI transparency

The federal government launched a consultation on July 23 seeking views on how to strengthen transparency for AI systems and for AI-generated content, a step toward disclosure rules under the AI for All national strategy. For Canadian organizations, it is the window to shape obligations before they are written — and a signal that labelling what a system is and what it produced is moving from good practice toward requirement.

Leave a Reply