Voluntary primary framework for incorporating trustworthiness into AI design, development, use, and evaluation; NIST notes that revision work is underway.
Finance AI evaluation scorecard for accounting and close workflows.
Score a finance AI system on evidence, authority, exceptions, security, and operating fit—not on the polish of one generated answer. This Sebastian-authored buyer worksheet deliberately separates a demonstration from a production control claim.
Evaluate one real workflow, not a generic promise.
Evaluate one bounded workflow with representative data and known failure cases. Record the system’s permitted reads and writes, the source data available to it, the human decision boundary, and the artifact retained for review. A buyer should be able to repeat the test and explain why the result passed or failed.
Use a zero-to-three score for each dimension: zero means absent or not evidenced; one means described but not demonstrated; two means demonstrated in a bounded test; three means implemented, monitored, and evidenced in the target operating environment. A high average cannot compensate for a zero on posting authority, tenant isolation, or evidence retention when those are material requirements.
Eight dimensions to score from 0–3
Current total: 0/24
- 01
Workflow boundary
Ask: Is the exact task, input, output, and prohibited action defined?
Retain: A testable scope statement and a negative test outside that scope.
- 02
Source provenance
Ask: Can a reviewer identify which source versions and parameters informed the result?
Retain: Citations or bound evidence with timestamps, scope, and completeness checks.
- 03
Accounting validation
Ask: Which deterministic checks run independently of model judgment?
Retain: Balance, required-field, period, duplicate, and policy tests with failure behavior.
- 04
Human authority
Ask: Where must a qualified person approve, reject, or modify the work?
Retain: Interrupt and permission tests showing the system cannot bypass the decision.
- 05
Exception behavior
Ask: Does the system stop safely on ambiguity, missing data, conflict, or unavailable tools?
Retain: Representative failure cases, escalation route, and idempotent retry behavior.
- 06
Security and tenancy
Ask: How are workspace scope, sensitive fields, private records, and credentials isolated?
Retain: Architecture, access tests, encryption scope, and operational responsibility map.
- 07
Evaluation and monitoring
Ask: Are quality, control failures, latency, and human overrides measured over time?
Retain: Versioned scenarios, expected outcomes, run history, and named monitoring owner.
- 08
Retained operating record
Ask: Can a later reviewer reconstruct the inputs, generated work, approvals, and disposition?
Retain: Durable state and append-only or versioned activity tied to the workflow record.
Score interpretation and stop rules
Set stop rules before the demo. Otherwise an impressive narrative can move the threshold after a material control fails.
| State / score | Minimum evidence | Decision rule |
|---|---|---|
| 0 — absent | No reliable evidence or the capability is outside current scope. | Stop if the dimension is material to the workflow. |
| 1 — described | Documentation or roadmap exists, but the buyer has not observed a bounded test. | Do not treat as implemented in the purchasing decision. |
| 2 — demonstrated | A representative bounded test passed with retained evidence. | Define production validation and monitoring before launch. |
| 3 — operating | Implemented in target conditions with monitoring, owner, and retrievable records. | Revalidate after material model, prompt, tool, or policy changes. |
Worked example: a demo that fails a material stop rule
This fictional scorecard shows why an average must not override a zero on a material control. It does not evaluate Sebastian or another named vendor.
| Field | Completed example | Interpretation |
|---|---|---|
| Workflow | AI-prepared payroll accrual | Bounded scenario with known source data and expected entry. |
| Observed scores | Boundary 2 · Provenance 2 · Validation 2 · Human authority 0 · Exceptions 1 · Security 2 · Monitoring 1 · Record 2 | Total 12/24, but the total is not the deciding signal. |
| Material stop rule | Human authority = 0 | The demo did not prove that preparation cannot bypass posting authorization. |
| Decision | Do not launch the write path | Keep the scenario non-posting until the approval boundary passes a negative test. |
Disclosed appendix: Sebastian self-assessment boundary
Sebastian publicly evidences dependency-aware close tasks, persistent review workbooks and activity, versioned source snapshots with content verification, durable encrypted agent state, an approval interrupt for agent journal creation, and balanced pending-review journal drafts.
The same scorecard should assign zero—not future credit—to claims that are not implemented. Today that includes generalized AP automation, generalized reconciliation, ERP write-back, accounting-period lock enforcement, and enforced segregation of duties.
Prospective design partners should run their own representative scenarios, preserve the test inputs and expected outcomes, and have finance, security, and system owners sign off on the production boundary. Sebastian’s public re-enactments are product explanations, not customer-performance evidence.
Standards context, with applicability boundaries.
Cross-sector companion resource addressing generative-AI risks across the lifecycle.
See the work up close.
Keep your ERP. Close and plan in one pane.
AI for accounting and FP&A on the ledger you already run. No migration. No rip and replace.