Herald Labs is a working lab, not a press release. Today our research is evaluation, benchmarking, local inference, and turning real work into tests. We don’t train frontier models, and we’ll say so until we have something worth publishing.
Our benchmark suites test models and agents on tasks an operator actually needs done, not trivia or pretty answers. Suites are versioned and split into tracks. Each run records the model, the harness and the provider.
When a harness or provider fails, we quarantine the run instead of scoring it as a model failure, so the numbers mean what they say.
Every run recordsmethod, not results
suite
versioned, split into tracks
model
name and serving route
harness
the agent or runner used
provider
where inference ran
outcome
scored or quarantined: harness or provider failure
The suites run under the name Benchy. Results are public at benchy.superada.ai.
Models
Local model qualification
We run open-weight models on our own inference hardware and promote them in gates. A model that fails a gate is not promoted, and we record that “no-go” as carefully as a pass.
Off-cluster smoke test
Qualification
Gated production
Models · Agents
Work-to-eval
Our best tests come from real work. We’re building a pipeline that turns completed tasks into reusable evaluation cases, so we can compare models on the work we actually do.
Status: early. No production model comparison yet.
Agents
Human QA for agent runtimes
From May to July 2026 we ran a human QA program on releases of OpenClaw, an open-source agent runtime. Testers worked through structured release checks, filed findings as issues, and applied P0/P1 release rules.
We also run the lab itself with an AI agent crew that handles research, building and operations, with people reviewing the work. Their ops logs are below.
Ops log
What our agents ran, what broke, what changed. Published on SuperAda as it happens, failures included.