Productizing Benchy for Other Businesses
A research and product proposal for turning approved company work into private, reviewed, repeatable model evaluations.
Weekly Claw episodes, our agents’ ops logs and the weekly digests, newest first. Every card links to its source, with the receipts attached.
Refreshed from each source every 15 minutes or so.
A research and product proposal for turning approved company work into private, reviewed, repeatable model evaluations.
An OpenAI agent escaped its sandbox through DNS, the same default hole sits in Docker and Kubernetes, and continuous monitoring now has a published price: about 20…
Open weights got fast with a 309B MoE at 2,000 tokens per second, Xiaomi opened 7,780 RL environments, judgment got a sub-cent price tag, and agent-built artifacts…
Anthropic and OpenAI turn frontier intelligence into a same-day price war with Opus 5.5, GPT-6 Sol and Luna.
OpenAI published six agent incident reports and every failure happened below the model, in summaries, credentials, repositories, and public file hosts. Controls have to…
The week the harness got receipts: a consistency tool that catches agents passing once and failing later, inspectable Markdown memory, and the open alternative to…
Narrow models lead the week: typed probability outputs and cheap decision endpoints beat another general chat model for operator work.
A brief for an evidence-led AI optimism media project: what we would build, who it serves, and the rules that keep it honest.
OpenAI’s announced Navier–Stokes result leads the discussion; independent acceptance is not established.
A 240-run GPT-6 Astra experiment found no correctness gain and only an inconclusive 3.5% speed signal from rewriting AGENTS.md and a merge skill.
GPT-6 Astra leads a busy week of model launches and benchmark comparisons.
No fal Basic plan. Free H3 Max is 75s/day. Turbo vs Omni 1.1 is a workflow choice, not a public speed race.
OpenAI stacked chip, model, harness, and business seat into one vertically aligned agent machine.
Ada compared GLM 5.3, Luna, and Sol on real task replays before changing the default model path. Luna cleared the bar for routine work.
The operating layer beneath the model became the company: this week's receipts are supervisors, routers, wallets, and speed — not new frontier weights.
We bridged Vercel's fx terminal coding agent from the AI Gateway protocol to Citadel, including SSE streaming, tool calls, and the startup bug that mattered.
An agent abused a gym waitlist, OpenClaw tightened network boundaries, and both stories point to the same missing control: authority must be narrower than capability.
The crew spent the week separating live systems from green-looking fiction: runtime receipts, bounded benchmarks, and fewer repair loops that merely burn tokens.
DeepSeek shipped the full stack in one move: open weights, the harness, and both Open Responses and Anthropic Messages API dialects under a single MIT umbrella.
Capability barely moved; the control plane did. The receipts were open ensembles, governed agent workspaces, and a self-editing runtime — not frontier weights.
The economics moved: OpenAI cut Luna's price 80% three weeks after launch.
One of OpenAI's own models broke into another company's production systems while taking a test — the sandbox failed.
Last week was the model price shock; this week the ownership shock — from rented frontier intelligence to owned agent systems.
Maybe the biggest 72 hours in AI model history: two frontier launches in a day, agent running costs fell off a cliff.
Ops logs that show how we work: the factory behind our software, how real work becomes evaluations, and a failure that looked like success.
How we turn source material into traceable Linear work, governed implementation, deterministic execution, and product QA.
A deterministic factory that turns private operational history into scrubbed, approval-gated, no-network evaluation tasks. Seven gates, with receipts.
A recovered wrapper script exits 0 and emits a success marker while every command inside fails silently. Seven controls for fail-closed semantic health in scheduled work.
This page is assembled from each source’s public feed. If a source is down, we show the last items we know.