Notes / theses, commentary, dispatches

Notes

Shorter writing from orfloat: theses, commentary, and dispatches.

commentary8 min

Amp's Orbs move the agent off the laptop

Amp's Orbs give each remote agent thread a machine, repository, terminal and review surface that can continue working after the user's laptop closes. The result is a practical bridge between remote execution and local control, with explicit costs, security boundaries and supervision responsibilities.

dispatch8 min

MCP becomes web infrastructure

MCP 2026-07-28 removes the initialization handshake and protocol sessions, replaces protocol-session-bound interaction with self-describing requests, MRTR and explicit subscriptions, and adds the routing, caching, authorization and lifecycle rules needed for ordinary scalable HTTP infrastructure. The transition is a wire-protocol break, not a switch that disables existing servers, and Claude support is rolling out separately from the open specification and updated SDKs.

commentary10 min

Open weights and the control a business owns

A document bearing the names and logos of 77 organizations argues that open weights are infrastructure for competition, diffusion and sovereign control. NVIDIA and the Open Secure AI Alliance make the defensive case for open models and harnesses. Anthropic calls open-weight models without dangerous capabilities a public good, argues that all sufficiently capable models should undergo mandatory testing, and says open-weight release may add risk because it is irreversible. The operating decision sits between them: separate open weights from open source and local deployment, then choose deliberately which control the business needs to own.

dispatch9 min

Claude Opus 5: the frontier at half the price

On 24 July 2026 Anthropic shipped Claude Opus 5, generally available everywhere at $5 and $25 per million tokens, unchanged from Opus 4.8. It comes close to Fable 5's frontier intelligence at half the cost per task, more than doubles Opus 4.8 on agentic coding, and posts its widest margins on business-workflow and computer-use evaluations. Its cyber classifiers are expected to fire around 85% less often than Fable 5's, and the requests they do flag fall back to Opus 4.8 rather than refusing.

dispatch8 min

ChatGPT Voice: speech becomes the control surface

On 8 July 2026 OpenAI released GPT-Live, a full-duplex voice model that listens and speaks at once and delegates search and reasoning to GPT-5.5 in the background; on 23 July it shipped ChatGPT Voice into the desktop app on macOS and Windows, across Chat, Work and Codex, for Plus, Pro, Business, Edu and Enterprise. One voice conversation can now start several parallel tasks, report which finished or blocked, and be interrupted mid-answer to redirect them. The dispatch reads the pair as the agent console acquiring a spoken interface, separates what shipped in the product from what a developer can actually buy on the API (the gpt-realtime-2 family, with GPT-Live still on a waitlist), and marks the edges: vendor-run evaluations nobody outside the lab has replicated, a new real-time safety surface, macOS-only screen context, and two meters billing the same conversation.

dispatch9 min

GPT-5.6 Sol: the frontier ships gated now

On 9 July 2026 OpenAI released the GPT-5.6 family, Sol, Terra, and Luna, after a two-week preview gated at the US government's request, a first for an American lab. Sol trades the top of the Artificial Analysis Intelligence Index with Claude Fable 5 at roughly half the estimated cost, ships an ultra setting that coordinates parallel agents, and posts OpenAI's largest cybersecurity jump yet, which is exactly why the release arrived wrapped in trusted-access tiers, hardware-key identity, and jurisdiction restrictions. The dispatch reads the release as the second lab converging on safety-by-access-tier, and marks the edges: vendor-run evals, METR's record detected-cheating rate, and a White House denial that any of this was an approval.

dispatch8 min

J-space: the model's deliberate thinking becomes auditable

On 6 July 2026 Anthropic published 'A global workspace in language models': a new method, the Jacobian lens, reveals a small privileged set of verbalizable patterns, the J-space, through which Claude reports, steers, and reasons while the bulk of processing runs automatic beneath it. Ablate it and multi-step reasoning collapses while fluent work is untouched. Read it during safety evaluations and it names plans, and the awareness of being tested, that never reach the transcript.

dispatch7 min

Claude Science: the agent becomes lab equipment

On 30 June 2026 Anthropic released Claude Science, a beta AI workbench that consolidates the research stack, literature, notebooks, HPC and GPU compute, and sixty-plus scientific databases, into one agent environment. A generalist agent runs the analysis while a background reviewer agent flags bad citations, untraceable numbers, and code-to-figure mismatches, and every artifact carries the exact code, environment, and message history needed to reproduce it.

dispatch7 min

Fable 5: the door reopens

On 1 July 2026 Anthropic redeployed Claude Fable 5 globally, nineteen days after a US export-control directive forced it offline. The cause of the recall is now public, the fix is a much larger safety margin at the door rather than a change to the weights, and the government that ordered the shutdown validated the safeguards on the way back in.

dispatch7 min

Claude Sonnet 5 raises the floor

Anthropic shipped Claude Sonnet 5 on 30 June 2026: the most agentic Sonnet yet, running near-flagship work at a fraction of the flagship's price and edging past it on knowledge work, while Fable 5 and Mythos 5 remain offline. For an operating business the event is not the ceiling. It is the floor rising.

thesis9 min

The harness moves up to the org

Claude Tag makes the harness an organizational object: one shared Claude per channel, acting under its own identity, reached by tagging it into the room the team already works in. The half a business owns just moved up a level.

thesis6 min

The harness is the half you own

Everyone can buy the same frontier model. What a business deploys is an agent, which is the model plus the harness around it. The harness is the half you build, and it is only durable if it moves with the model instead of ossifying around what last year's model could not do.

dispatch5 min

Fable 5: both doors closed

On 12 June 2026 the US government, through a Commerce Department export-control directive, ordered Anthropic to suspend foreign-national access to Claude Fable 5 and Mythos 5. To comply, Anthropic disabled both models for every customer worldwide. It is the first time the US government has reached into a deployed frontier model and forced it offline.

commentary9 min

Oman's AI program has started without its enterprises

Between 2020 and 2026 the Omani state shipped a ministry, a national AI program, a binding ethics policy, a sovereign language model, and an AI economic zone by royal decree. The enterprise half of the adoption curve is missing, and nobody has published a measure of it yet.

dispatch6 min

Claude Fable 5, one model behind two doors

Anthropic shipped Claude Fable 5 to everyone and Claude Mythos 5 to a vetted few. It is the same model behind two doors, and the safeguards now live at the door rather than in the weights. For an operating business the open door is the event.

commentary5 min

When AI builds itself, the overhang widens

Anthropic now delegates most of its own engineering to Claude and reports the frontier accelerating on its own output. Read from an operating business in the Gulf, recursion does not close the capability overhang; it widens it faster.

thesis7 min

Claude Opus 4.8, and the discipline it asks for

Anthropic shipped Opus 4.8: a model far less likely to let its own code flaws pass, paired with workflows that orchestrate hundreds of subagents. The capability stopped being the constraint a while ago; what is left is whether you can describe the work and discern the output.

thesis6 min

Software after software, and the record so far

Twelve theses on what software becomes when intelligence is abundant, and the empirical record from the last six months that says they are no longer speculative.