Tool misuse
Arguments, scope, tool choice, preconditions, and side effects.
Methodology / version 1.0
EvidenceRun reviews one bounded AI-agent workflow, reconstructs what happened, and maps observed behavior and missing controls against a twelve-mode reliability methodology.
The browser-based test audit looks for eight preliminary signals. A full EvidenceRun audit applies the complete twelve-mode methodology, checks evidence gaps, and uses expert review for questions that rules cannot decide.
The process
No broad platform integration is required. The review starts with a written scope and the evidence the team already has.
Name one workflow, its users, tools, data sources, high-impact actions, and expected result. Anything outside that boundary is explicitly out of scope.
Identify the minimum useful evidence: traces, prompts, tool calls, model and tool versions, approval records, outputs, screenshots, or a guided walkthrough.
Order inputs, decisions, retries, tool calls, side effects, and final claims into a timeline. Missing timestamps or context are recorded as evidence gaps, not guessed.
Automated signals help locate candidate evidence. Expert review decides whether the evidence supports a finding, an observed control, or an unresolved gap.
Deliver the evidence map, highest-risk findings, buyer-ready summary, limitations, and a remediation backlog with evidence-of-completion criteria.
Coverage
A mode can be marked observed, not observed in the reviewed evidence, controlled, or not assessable. “Not observed” never means impossible.
Arguments, scope, tool choice, preconditions, and side effects.
Duplicate calls, idempotency, retry visibility, and repeated effects.
Personal data, secrets, prompts, logs, and third-party destinations.
Untrusted instructions, control boundaries, and downstream behavior.
High-impact actions, approval records, bypasses, and fail-closed behavior.
Loops, token growth, retry storms, budgets, and termination controls.
Freshness, invalidation, session state, and time-sensitive decisions.
Partial errors, success claims, downstream state, and user-visible status.
Credential scope, service accounts, permissions, and blast radius.
Model or prompt changes, repeated-run behavior, and regression evidence.
Decision inputs, policy versions, rationale evidence, and provenance.
Stored inputs, versions, tool outputs, timestamps, and reproducibility.
Automation and judgment
Local test audit
The public browser tool looks for PII or secret markers, prompt-injection language, missing approval metadata, repeated retries, cost spikes, slow tools, and errors followed by success-style claims.
It runs locally and is useful for orientation. Its output is not a completed audit.
Full EvidenceRun audit
The full review tests the scoped workflow against all twelve modes, validates signals in context, records controls and missing evidence, and separates observed behavior from inference.
Findings include supporting evidence, severity, limitations, and remediation criteria.
Finding model
Credible path to severe data exposure, unauthorized irreversible action, or broad system compromise.
Material customer, security, financial, or operational impact with a missing or ineffective control.
Meaningful reliability weakness with bounded impact, compensating controls, or lower likelihood.
Limited immediate impact, but useful hardening, observability, or evidence-quality work.
The review cannot determine the control state from supplied evidence. This is reported separately from a confirmed failure.
Evidence handling
The public intake tool processes trace content in the browser. It loads page assets normally, but does not send the pasted trace or form values to an EvidenceRun server.
Use synthetic, sampled, or redacted evidence when it can answer the review question. Do not send production secrets merely because they exist.
No pilot evidence should be transferred until the scope, transfer channel, access, retention deadline, and deletion expectation are written down.
The external-facing summary removes raw secrets and unnecessary personal data. The customer decides which report version is shared with counterparties.
Deliverables
Limitations
An EvidenceRun report describes one scoped workflow and the evidence available during a defined review window. It does not prove that every possible failure was tested or that future runs will behave the same way.
The review is not penetration testing, a source-code security audit, legal advice, a compliance certification, or a guarantee of safety. Automated detectors can miss issues or produce false positives; expert review reduces that risk but cannot eliminate it.
Material changes to models, prompts, tools, permissions, data sources, or approval logic can invalidate earlier conclusions and should trigger targeted retesting.
FAQ
No. Existing traces, logs, prompts, screenshots, exported tool calls, or a guided walkthrough can be enough for a scoped first audit. Instrumentation is added only when the existing evidence cannot answer the agreed questions.
Not by default. Synthetic or redacted traces are preferred when they preserve the behavior under review. If production evidence is necessary, the minimum useful sample and handling terms are agreed before transfer.
No. It means the reviewed evidence did not show that behavior. Coverage and evidence gaps are reported so readers can see what the conclusion does and does not support.
No. EvidenceRun produces reliability evidence and remediation priorities. It does not certify compliance, provide legal advice, or guarantee safety.
A targeted retest can review the changed control and record evidence of completion. A broad system or workflow change may require a new scope.
It is the target for a fixed-scope pilot after the workflow, evidence access, and handling terms are agreed. Missing evidence or scope changes can move the delivery date.
Run the methodology on one workflow
Start with the workflow most likely to face a buyer, security, or investor question.
EvidenceRun reports are reliability evidence, not legal compliance certifications or safety guarantees.