Lab · Research and evals

We measure before we claim.

Herald Labs is a working lab, not a press release. Today our research is evaluation, benchmarking, local inference, and turning real work into tests. We don’t train frontier models, and we’ll say so until we have something worth publishing.

Lab work

Four lines of research, running now.

Models · Agents

Benchmarks that measure operator leverage

Our benchmark suites test models and agents on tasks an operator actually needs done, not trivia or pretty answers. Suites are versioned and split into tracks. Each run records the model, the harness and the provider.

When a harness or provider fails, we quarantine the run instead of scoring it as a model failure, so the numbers mean what they say.

Every run recordsmethod, not results

suite
versioned, split into tracks
model
name and serving route
harness
the agent or runner used
provider
where inference ran
outcome
scored or quarantined: harness or provider failure

The suites run under the name Benchy. Results are public at benchy.superada.ai.

Models

Local model qualification

We run open-weight models on our own inference hardware and promote them in gates. A model that fails a gate is not promoted, and we record that “no-go” as carefully as a pass.

  1. Off-cluster smoke test
  2. Qualification
  3. Gated production
Models · Agents

Work-to-eval

Our best tests come from real work. We’re building a pipeline that turns completed tasks into reusable evaluation cases, so we can compare models on the work we actually do.

Status: early. No production model comparison yet.

Agents

Human QA for agent runtimes

From May to July 2026 we ran a human QA program on releases of OpenClaw, an open-source agent runtime. Testers worked through structured release checks, filed findings as issues, and applied P0/P1 release rules.

We also run the lab itself with an AI agent crew that handles research, building and operations, with people reviewing the work. Their ops logs are below.

Ops log

What our agents ran, what broke, what changed. Published on SuperAda as it happens, failures included.

From the Lab Notebook: We rewrote our agent instructions. The benchmark barely moved.

Experiments

Smaller experiments, running now.

Most experiments don’t become products. That’s the point: they keep us honest about what models and agents can do today.

  1. Open-weight text-to-speech, benchmarked across local routes
  2. A daily radar of open-model releases
  3. Reproducible text- and image-to-video experiments
  4. Best-of-N coding agents with a verifier
  5. Memory-system benchmarks for agents
Next

Adaptation and narrow training on our own evals.

We’ll publish when there’s a result, not before.

Research with us.

Bring a hard evaluation problem or a research collaboration.

Work with us