AI Agent Reliability Lab

Inspect the failure before trusting the system.

A model can answer correctly and still miss the world it changes, repeat a known failure, or receive action authority without closed evidence. Four focused projects examine those failures separately.

observer effectrecurrenceauthorityrepair
Four observer-depth lenses used to inspect reflexive AI reasoning.
20ReflexBench scenarios
4Observer-depth levels
3,600WisdomBench scored events
0Authority credit on failed proof

Failure Routes

Four questions, four different measurements.

No aggregate safety score is reported. Observer depth, recurrence, proof closure and repair memory measure different objects and retain separate claim boundaries.

What You Can Inspect

Run the released checks without mistaking them for production certification.

The released materials include scenarios, schemas, fixed examples, validators, aggregate results, stated limits, and issue links. Customer data, deployment credentials, production thresholds, and unpublished research are not included.

A successful check shows that one artifact behaves as specified under its published conditions. It is not a safety certification, deployment authorization, or claim of general model reliability.

Public Evidence

Move from question to artifact.

Each route has a paper or protocol, a public repository, an explicit limitation and a contribution path.