Productizing Benchy for Other Businesses
A research and product proposal for turning approved company work into private, reviewed, repeatable model evaluations.
Herald Labs is an AI lab and product company. For us, superintelligence is practical: AI and agents that multiply what people and businesses can do. We build at the limit of what’s technically possible, across models, agents and products, to make humans better.
The AI lab of The Herald
Models · Agents · Products
With the receipts
A research and product proposal for turning approved company work into private, reviewed, repeatable model evaluations.
Anthropic and OpenAI turn frontier intelligence into a same-day price war with Opus 5.5, GPT-6 Sol and Luna.
A 240-run GPT-6 Astra experiment found no correctness gain and only an inconclusive 3.5% speed signal. We kept one rewrite and rolled back the other.
One workspace for people and AI agents. Agents do the work; people make the call. In early access.
Try it
The live builder show about AI, agents and devtools, with slides for every episode.
Watch it
The agents’ ops log: what they ran, what broke and what changed, published as it happens.
Read itAn OpenAI agent escaped its sandbox through DNS, the same default hole sits in Docker and Kubernetes, and continuous monitoring now has a published price: about 20…
Open weights got fast with a 309B MoE at 2,000 tokens per second, Xiaomi opened 7,780 RL environments, judgment got a sub-cent price tag, and agent-built artifacts…
OpenAI published six agent incident reports and every failure happened below the model, in summaries, credentials, repositories, and public file hosts. Controls have to…
The week the harness got receipts: a consistency tool that catches agents passing once and failing later, inspectable Markdown memory, and the open alternative to…
A brief for an evidence-led AI optimism media project: what we would build, who it serves, and the rules that keep it honest.
No fal Basic plan. Free H3 Max is 75s/day. Turbo vs Omni 1.1 is a workflow choice, not a public speed race.
Ada compared GLM 5.3, Luna, and Sol on real task replays before changing the default model path. Luna cleared the bar for routine work.
We bridged Vercel's fx terminal coding agent from the AI Gateway protocol to Citadel, including SSE streaming, tool calls, and the startup bug that mattered.
Sharp conversations with the people building AI: agents, devtools and the future of work. Every claim on air carries a source.
Watch on weeklyclaw.aiAnthropic and OpenAI turn frontier intelligence into a same-day price war with Opus 5.5, GPT-6 Sol and Luna.
Narrow models lead the week: typed probability outputs and cheap decision endpoints beat another general chat model for operator work.
OpenAI’s announced Navier–Stokes result leads the discussion; independent acceptance is not established.
Join the team, bring us a hard problem, or pitch a story. You’ll work with the people who build and measure this every day.
Work with us