Build the
evidence loop.
A public evaluation and citation policy for systems that help people make evidence-sensitive decisions.
Research is a lifecycle, not a leaderboard.
Our working model follows four questions: what risk is in scope, what evidence would change a decision, how should uncertainty be recorded, and how will the process be maintained after launch?
That approach is informed by the NIST AI Risk Management Framework, which organizes risk work around Govern, Map, Measure, and Manage. NIST’s guidance is a useful reminder that measurement should resemble deployment conditions, not only a tidy laboratory.
Our evaluation policy
We define the task before selecting a metric. A lead-finding system, a source-traceability system, and a decision-support system have different denominators and different costs of error. Evaluation sets should document scope, time window, language, claim type, exclusions, and the reference process.
We report abstentions, disagreements, corrections, and cases where a citation is relevant but insufficient. We inspect performance across source quality, topic, time delay, and context—not only an aggregate. When a retrieval source, model, prompt, or threshold changes, the evaluation is a new snapshot rather than a permanent product property.
Citation policy
A citation must be identifiable, accessible where possible, and connected to the specific part of the claim it supports. We prefer primary records for primary assertions, then transparent reporting and specialist context. A collection of copied pages is not a set of independent confirmations.
For synthetic media and content history, we distinguish provenance assertions, labels, detection signals, testing, auditing, and maintenance. The C2PA specification can bind signed assertions to content, but it does not establish that a depicted event is true; missing credentials do not establish manipulation. The NIST AI RMF and OECD AI Principles both support a contextual, accountable approach.
Why people remain in the loop
The OECD’s work on information integrity emphasizes transparent, plural sources and hybrid human-plus-technical approaches. No single technical method resolves misinformation in every context. Our systems should expose the path to a decision, make correction possible, and leave room for expertise.