Skip to content
LiveNext 12:33:16
Today's five/N°6 · Sunday, October 11, 2026/Filed 04:48 ET

AI news today: Microsoft's cheap decision model, Nadella's emergency brake for AI

Agents got cheaper to run and harder to trust on the same weekend. The fix for both is the same: put the checks outside the model.

The five · ranked by consequence

Permalink ↗
  1. Microsoft's Decision-1 makes the yes/no call cheap

    Microsoft released Decision-1, a Qwen-based model that scores fixed choices, at $0.042 per million input tokens with free output. The benchmarks are Microsoft's own. The price is the reason to move your agent's small calls off big models.

    Why it matters If your agent sends routing, triage and pass/fail checks to a frontier model, you are paying frontier prices for one-word answers. A cheap scorer with a probability you can threshold is worth testing on those calls this week.

    SRCMicrosoft

  2. Nadella says assume AI models are compromised

    Microsoft CEO Satya Nadella called on Saturday for an AI "emergency brake": controls outside the model, tamper-proof logs and a human who can stop a task mid-run. Treat it as a checklist for your own agents, not a press line.

    Why it matters If the model is assumed compromised, safety is no longer something you buy from a lab. It is the harness you build around it. Check today whether your agents can edit their own permissions or logs, and whether a person can stop them mid-task.

    SRCTechCrunch

  3. Claude and Codex agents decompiled a shooter, 83% exact

    A developer says Claude and Codex agents rebuilt a popular first-person shooter as C++ over about three months, with 83% of functions byte-exact. The game is the headline. The locked test harness that kept the agents honest is the part to copy.

    Why it matters Long agent jobs work when "done" is a check the agent cannot edit. If you run multi-agent coding tasks, build that check first. If you ship compiled software, assume reconstructing your source just got much cheaper.

    SRCmomo5502.com

  4. Fake Claude installer ads target Mac developers

    Push Security says a Google ad for "claude mac" showed bing.com, then routed users to a fake Claude page whose copy button swaps in a malicious command. If you install AI tools from a search ad, stop.

    Why it matters AI coding tools install with one pasted terminal command, and the people pasting them hold source code and cloud keys. Type the vendor's address yourself, paste install commands into a text editor first, and tell your team today.

    SRCPush Security

  5. Nvidia is in talks to buy Reflection AI, FT reports

    Nvidia is in talks to acquire open-weight model startup Reflection AI or invest more in it, the Financial Times reported, days after Reflection announced its 501B-parameter Beam model. Nothing changes for you until Beam's weights ship under Apache 2.0.

    Why it matters An open model owned by the dominant chip supplier has every reason to favor that supplier's hardware. If you plan to build on Beam, the Apache 2.0 weights are what you can count on. Download them when they ship, whoever owns the company.

    SRCBloomberg

Prompt of the day

Audit your AI agent for an emergency brake

› Act as a blunt reviewer of AI agent safety and cost. I will describe one AI agent or automation I run. Audit it against these five checks and give each a PASS, FAIL or UNKNOWN with one line of reasoning:

1. Separation: can the agent change its own instructions, permissions, tools or logs? It should not be able to.
2. Outside controls: which limits (network, files, spending, accounts) are enforced outside the model, and which rely on the model behaving?
3. Evidence: does every action that touches the outside world (sending, submitting, buying, writing to a system) leave a record a person can read later and the agent cannot edit?
4. Brake: can a named person pause or stop a task mid-step without redeploying? Who, and how?
5. Cheap decisions: list every step where the agent is really making a yes/no or pick-one call (routing, triage, pass/fail checks). Flag which could go to a smaller, cheaper model with a confidence threshold, and which need a human below that threshold.

Then give me the three fixes that cut the most risk for the least work, in order, each with a first step I can do this week. Ask me questions first if my description leaves a check UNKNOWN.

My agent: [describe what it does, which model it uses, what tools and accounts it can reach, and who watches it]

Works in any AI chat. In Magai, run it against several models at once.

Every day, on the record.

Each square is an edition. Click any day to read the five that mattered then.

Editions
6
Stories ranked
30
Topics tracked
10
Terms defined
24
AprMayJunJulAugSepOct
EditionNo editionTodayFull archive →

Speak fluent AI.

Every term our stories use, defined in plain English and linked wherever it appears.

The glossary →