Controller resource · Vendor evaluation

Finance AI evaluation scorecard for accounting and close workflows.

Score a finance AI system on evidence, authority, exceptions, security, and operating fit—not on the polish of one generated answer. This Sebastian-authored buyer worksheet deliberately separates a demonstration from a production control claim.

Prepared by Sebastian Product & EngineeringUpdated July 20, 2026Educational worksheet · not audit advice
How to use it

Evaluate one real workflow, not a generic promise.

Evaluate one bounded workflow with representative data and known failure cases. Record the system’s permitted reads and writes, the source data available to it, the human decision boundary, and the artifact retained for review. A buyer should be able to repeat the test and explain why the result passed or failed.

Use a zero-to-three score for each dimension: zero means absent or not evidenced; one means described but not demonstrated; two means demonstrated in a bounded test; three means implemented, monitored, and evidenced in the target operating environment. A high average cannot compensate for a zero on posting authority, tenant isolation, or evidence retention when those are material requirements.

Original controller worksheet

Eight dimensions to score from 0–3

Current total: 0/24

  1. 01

    Workflow boundary

    Ask: Is the exact task, input, output, and prohibited action defined?

    Retain: A testable scope statement and a negative test outside that scope.

  2. 02

    Source provenance

    Ask: Can a reviewer identify which source versions and parameters informed the result?

    Retain: Citations or bound evidence with timestamps, scope, and completeness checks.

  3. 03

    Accounting validation

    Ask: Which deterministic checks run independently of model judgment?

    Retain: Balance, required-field, period, duplicate, and policy tests with failure behavior.

  4. 04

    Human authority

    Ask: Where must a qualified person approve, reject, or modify the work?

    Retain: Interrupt and permission tests showing the system cannot bypass the decision.

  5. 05

    Exception behavior

    Ask: Does the system stop safely on ambiguity, missing data, conflict, or unavailable tools?

    Retain: Representative failure cases, escalation route, and idempotent retry behavior.

  6. 06

    Security and tenancy

    Ask: How are workspace scope, sensitive fields, private records, and credentials isolated?

    Retain: Architecture, access tests, encryption scope, and operational responsibility map.

  7. 07

    Evaluation and monitoring

    Ask: Are quality, control failures, latency, and human overrides measured over time?

    Retain: Versioned scenarios, expected outcomes, run history, and named monitoring owner.

  8. 08

    Retained operating record

    Ask: Can a later reviewer reconstruct the inputs, generated work, approvals, and disposition?

    Retain: Durable state and append-only or versioned activity tied to the workflow record.

Decision tool

Score interpretation and stop rules

Set stop rules before the demo. Otherwise an impressive narrative can move the threshold after a material control fails.

State / scoreMinimum evidenceDecision rule
0 — absentNo reliable evidence or the capability is outside current scope.Stop if the dimension is material to the workflow.
1 — describedDocumentation or roadmap exists, but the buyer has not observed a bounded test.Do not treat as implemented in the purchasing decision.
2 — demonstratedA representative bounded test passed with retained evidence.Define production validation and monitoring before launch.
3 — operatingImplemented in target conditions with monitoring, owner, and retrievable records.Revalidate after material model, prompt, tool, or policy changes.
Completed teaching example

Worked example: a demo that fails a material stop rule

This fictional scorecard shows why an average must not override a zero on a material control. It does not evaluate Sebastian or another named vendor.

FieldCompleted exampleInterpretation
WorkflowAI-prepared payroll accrualBounded scenario with known source data and expected entry.
Observed scoresBoundary 2 · Provenance 2 · Validation 2 · Human authority 0 · Exceptions 1 · Security 2 · Monitoring 1 · Record 2Total 12/24, but the total is not the deciding signal.
Material stop ruleHuman authority = 0The demo did not prove that preparation cannot bypass posting authorization.
DecisionDo not launch the write pathKeep the scenario non-posting until the approval boundary passes a negative test.
Product boundary

Disclosed appendix: Sebastian self-assessment boundary

Sebastian publicly evidences dependency-aware close tasks, persistent review workbooks and activity, versioned source snapshots with content verification, durable encrypted agent state, an approval interrupt for agent journal creation, and balanced pending-review journal drafts.

The same scorecard should assign zero—not future credit—to claims that are not implemented. Today that includes generalized AP automation, generalized reconciliation, ERP write-back, accounting-period lock enforcement, and enforced segregation of duties.

Prospective design partners should run their own representative scenarios, preserve the test inputs and expected outcomes, and have finance, security, and system owners sign off on the production boundary. Sebastian’s public re-enactments are product explanations, not customer-performance evidence.

Primary references

Standards context, with applicability boundaries.

Go deeper

See the work up close.

Sebastian accounting

See current accounting workflows and explicit implementation boundaries.

Finance data architecture

Inspect tenancy, encryption, evidence, and durable agent state.

Security and controls

Review exact protections, approval boundaries, and non-claims.

Keep your ERP. Close and plan in one pane.

AI for accounting and FP&A on the ledger you already run. No migration. No rip and replace.

Agents at work · Matching an invoice to its PO line…