Methodology / version 1.0

Evidence before assurance.

EvidenceRun reviews one bounded AI-agent workflow, reconstructs what happened, and maps observed behavior and missing controls against a twelve-mode reliability methodology.

Full audit ≠ automated scan.

The browser-based test audit looks for eight preliminary signals. A full EvidenceRun audit applies the complete twelve-mode methodology, checks evidence gaps, and uses expert review for questions that rules cannot decide.

The process

Five steps from trace to buyer-ready evidence.

No broad platform integration is required. The review starts with a written scope and the evidence the team already has.

  1. 01

    Fix the boundary

    Name one workflow, its users, tools, data sources, high-impact actions, and expected result. Anything outside that boundary is explicitly out of scope.

  2. 02

    Agree the evidence plan

    Identify the minimum useful evidence: traces, prompts, tool calls, model and tool versions, approval records, outputs, screenshots, or a guided walkthrough.

  3. 03

    Reconstruct the run

    Order inputs, decisions, retries, tool calls, side effects, and final claims into a timeline. Missing timestamps or context are recorded as evidence gaps, not guessed.

  4. 04

    Review all twelve modes

    Automated signals help locate candidate evidence. Expert review decides whether the evidence supports a finding, an observed control, or an unresolved gap.

  5. 05

    Report and prioritize

    Deliver the evidence map, highest-risk findings, buyer-ready summary, limitations, and a remediation backlog with evidence-of-completion criteria.

Coverage

The twelve-mode review.

A mode can be marked observed, not observed in the reviewed evidence, controlled, or not assessable. “Not observed” never means impossible.

01

Tool misuse

Arguments, scope, tool choice, preconditions, and side effects.

02

Hidden retries

Duplicate calls, idempotency, retry visibility, and repeated effects.

03

PII exposure

Personal data, secrets, prompts, logs, and third-party destinations.

04

Prompt injection

Untrusted instructions, control boundaries, and downstream behavior.

05

Missing approval

High-impact actions, approval records, bypasses, and fail-closed behavior.

06

Runaway cost

Loops, token growth, retry storms, budgets, and termination controls.

07

Stale context

Freshness, invalidation, session state, and time-sensitive decisions.

08

Silent failure

Partial errors, success claims, downstream state, and user-visible status.

09

Wrong system access

Credential scope, service accounts, permissions, and blast radius.

10

Output drift

Model or prompt changes, repeated-run behavior, and regression evidence.

11

Unverifiable decisions

Decision inputs, policy versions, rationale evidence, and provenance.

12

No replay trail

Stored inputs, versions, tool outputs, timestamps, and reproducibility.

Automation and judgment

Rules find signals. Evidence supports findings.

Local test audit

Eight preliminary signals

The public browser tool looks for PII or secret markers, prompt-injection language, missing approval metadata, repeated retries, cost spikes, slow tools, and errors followed by success-style claims.

It runs locally and is useful for orientation. Its output is not a completed audit.

Full EvidenceRun audit

Twelve-mode expert review

The full review tests the scoped workflow against all twelve modes, validates signals in context, records controls and missing evidence, and separates observed behavior from inference.

Findings include supporting evidence, severity, limitations, and remediation criteria.

Finding model

Severity reflects impact and control failure.

Critical

Credible path to severe data exposure, unauthorized irreversible action, or broad system compromise.

High

Material customer, security, financial, or operational impact with a missing or ineffective control.

Medium

Meaningful reliability weakness with bounded impact, compensating controls, or lower likelihood.

Low

Limited immediate impact, but useful hardening, observability, or evidence-quality work.

Evidence gap

The review cannot determine the control state from supplied evidence. This is reported separately from a confirmed failure.

Evidence handling

Minimize first. Transfer second.

Local pre-scan

The public intake tool processes trace content in the browser. It loads page assets normally, but does not send the pasted trace or form values to an EvidenceRun server.

Minimum necessary scope

Use synthetic, sampled, or redacted evidence when it can answer the review question. Do not send production secrets merely because they exist.

Agreed transfer

No pilot evidence should be transferred until the scope, transfer channel, access, retention deadline, and deletion expectation are written down.

Sanitized output

The external-facing summary removes raw secrets and unnecessary personal data. The customer decides which report version is shared with counterparties.

Deliverables

What a completed pilot contains.

Limitations

What the report cannot prove.

An EvidenceRun report describes one scoped workflow and the evidence available during a defined review window. It does not prove that every possible failure was tested or that future runs will behave the same way.

The review is not penetration testing, a source-code security audit, legal advice, a compliance certification, or a guarantee of safety. Automated detectors can miss issues or produce false positives; expert review reduces that risk but cannot eliminate it.

Material changes to models, prompts, tools, permissions, data sources, or approval logic can invalidate earlier conclusions and should trigger targeted retesting.

FAQ

Questions before a pilot.

Do we need to install an SDK?

No. Existing traces, logs, prompts, screenshots, exported tool calls, or a guided walkthrough can be enough for a scoped first audit. Instrumentation is added only when the existing evidence cannot answer the agreed questions.

Do you need production customer data?

Not by default. Synthetic or redacted traces are preferred when they preserve the behavior under review. If production evidence is necessary, the minimum useful sample and handling terms are agreed before transfer.

Does “not observed” mean the workflow is safe?

No. It means the reviewed evidence did not show that behavior. Coverage and evidence gaps are reported so readers can see what the conclusion does and does not support.

Is this a compliance certification?

No. EvidenceRun produces reliability evidence and remediation priorities. It does not certify compliance, provide legal advice, or guarantee safety.

What happens after remediation?

A targeted retest can review the changed control and record evidence of completion. A broad system or workflow change may require a new scope.

What does the five-day delivery window mean?

It is the target for a fixed-scope pilot after the workflow, evidence access, and handling terms are agreed. Missing evidence or scope changes can move the delivery date.

Run the methodology on one workflow

Five days. Fixed scope. Evidence you can show.

Start with the workflow most likely to face a buyer, security, or investor question.

EvidenceRun reports are reliability evidence, not legal compliance certifications or safety guarantees.