Journal6 entries2026

Field notes. From the reliability lab.

Research, product notes and field lessons on catching agents that go rogue - evaluation, trace analysis and failure detection, written by the team building it.

Latest001Aug 14, 2026Research

Reproducing LLM-as-a-Verifier with open-source models on Terminal-Bench 2.0

We reproduce the verification results from LLM-as-a-Verifier with GLM-5.2, testing continuous score distributions, repeated evaluations, and criteria decomposition on Terminal-Bench 2.0.

Robert Hommes
Robert Hommes
Read entry
Showing 6 of 6

Reading up on agent reliability?

See what Moyai finds in your traces. First agent scan is free.

Upload traces