ONGOING

Measuring where automation succeeds and fails.

We evaluate extraction accuracy, hallucination rate, provenance accuracy, labor savings, and the share of assertions that require correction.

/work/ai-research-evaluationCultureOS Labs · 2026

MEASURES

Accuracy is not one number.

01

Extraction

Was the source read correctly?

02

Reasoning

Does the conclusion follow from the evidence?

03

Review

What did a scholar correct, reject, or investigate?

DISCIPLINE

Failure cases are results.

We publish where automation failed, where uncertainty remained, and where human review changed the outcome.

START HERE

What is your institution carrying?

We start with the one workflow costing you the most.