AXIGNAL
Get started
Topic · Evidence and knowledge boundaries

Model confidence does not measure evidence quality

Models may return a label and confidence score that look precise. That output describes how a provider evaluated a question, not how much real-world support exists for a claim about an organization.

In brief

AXIGNAL treats provider confidence as metadata about a replaceable evaluation. Evidence quality requires reviewing sources, provenance, currentness, contradictions, coverage and state. A model probability does not become a truth percentage, sale probability or universal epistemic meter.

Ask what the probability represents

A probability may describe which of several typed choices best fits the state shown to an evaluator. It does not necessarily represent observed frequency, calibrated accuracy or the likelihood of a commercial outcome. Without an appropriate contract and validation, reading the number as general certainty adds a meaning the provider did not promise.

Assess evidence separately

Support for a claim depends on the material and its conditions: who published it, which subject it describes, what period it covers and whether independent sources support it. Contradictions and missing information matter too. This review is multidimensional and may lead to more observation, research, POTENTIAL, UNKNOWN or admission under the relevant contract.

Retain evaluations for replay without granting authority

A structured evaluation may retain provider, model, version, contract, choices, fingerprints, distribution and observable conditions. This supports comparison across runs or reassessment under another policy. Recording the output makes it inspectable but does not make it a source of truth. The evaluator remains replaceable and non-authoritative.

Explain uncertainty in useful language

The interface should say what is supported, what is interpreted and what is missing. A precise-looking figure can suggest that AXIGNAL knows an objective probability even when sources are incomplete or a model is not calibrated for the task. If the question is not answerable, the system should investigate or abstain instead of using a high score to conceal the gap.

Conceptual example

Hypothetical: an evaluator assigns 0.86 to “relevant capability” among several categories. AXIGNAL retains that distribution and version, while showing that the only source is an old description and current availability is unsupported. The score does not mean an 86% chance the company can fulfill a contract.

Scope and limits

Provider scores are not interchangeable across models, contracts or versions and should not be surfaced without context. No universal threshold turns model confidence into admitted truth. Interpretation policy should be versioned and explainable, and remain separate from EvidenceAdmission and commercial outcomes.

Editorial basis

This page explains product doctrine and boundaries. It does not demonstrate source coverage or observed outcomes.

Read the product model

Understand AI evaluations

Explore AXIGNAL