Evidence-Gated Engine Layer

FinOps Engine Thinking Flow

Turning untrusted cloud-finance evidence into a verified FinOps diagnosis and a confidence-bounded path to action. One governed, evidence-gated agentic workflow. Six parallel forensic domains. Separate maturity and anti-pattern streams. Deterministic validation, tactic permissioning, independent fact-checking, and visible GO / WARN / BLOCK boundaries.
Interactive Version
30 Maturity + 30 Anti-Pattern Criteria
180 Evidence Questions
How to read this visual: the default layer shows the complete control architecture. Each clickable card then reveals the key question, reasoning logic, grounding basis, reliability / hallucination control, and an illustrative thinking sample. The purpose is to show how the Engine decomposes expert FinOps reasoning without exposing prompts.
AI reasoning Deterministic control Gate / validation Knowledge / support lane Publication boundary

1. Source intake, parsing & safety boundary

The Engine begins by turning heterogeneous source material into a safe, attributable assessment pack before any FinOps judgment is allowed.

Evidence boundary first
Critical trust idea: customer documents are untrusted evidence. Reference Knowledge Base material can shape interpretation and tactics, but it cannot become proof of the assessed organization’s current state.
PDF
Open logic
Input preparation

Multi-format local parsing

Normalizes PDFs, HTML, tables, JSON, screenshots, and other uploaded artifacts before model analysis.

Local extractionPage markersTable rows
Key question
What usable text, visual evidence, table structure, and source location can be extracted from each file?
Reasoning logic
Create machine-readable source material while preserving document, page, row, and visual context for later evidence checking.
Grounding basis
Local PDF extraction, HTML sanitation, table sampling, JSON normalization, image inputs, and parse-quality metadata.
Reliability control
The model is not used as the primary parser. Sparse or ambiguous extraction remains a warning instead of being silently treated as complete evidence.
Illustrative thinking sample“A policy statement on page 12 and a cost table on row 8 must remain independently traceable; they cannot become one blended contextual impression.”
DLP
Open logic
Safety control

DLP & sensitive-data review

Scans the source pack for secrets and sensitive identifiers before model calls and report generation.

SecretsRedactionDistributed scan
Key question
Can the material enter the pipeline without exposing credentials, restricted identifiers, or sensitive financial details?
Reasoning logic
Run deterministic and model-assisted safety review across the source registry, then redact or block unsafe material before deeper analysis.
Grounding basis
DLP patterns, registry-wide sampling, pre-flight review packets, and privacy-oriented report sanitation.
Reliability control
Financially sensitive source values can support diagnosis internally while being generalized or omitted from report-facing prose.
Illustrative thinking sample“Use the existence of a budget-control gap as evidence; do not echo account numbers, customer names, access tokens, or unnecessary exact spend values.”
PQ
Open logic
Input quality

Parse quality & visual evidence

Identifies sparse pages, weak extraction, and dashboard-like content that may require visual evidence handling.

Sparse-page warningVisual sourceQuality metadata
Key question
Does the extracted material adequately represent what a human reviewer can actually see in the file?
Reasoning logic
Keep selected visual pages or screenshots alongside extracted text so architecture diagrams, dashboards, and tables are not lost.
Grounding basis
Parse warnings, page density, image inputs, visual-page references, and source registry metadata.
Reliability control
A visually rich but text-poor page cannot be treated as “no evidence” solely because text extraction was sparse.
Illustrative thinking sample“A cost-allocation dashboard may prove role-based visibility even when its labels are embedded in an image rather than extractable text.”
KB
Open logic
Clean-room rule

Customer evidence vs. reference KB

Separates witness evidence from methodology, definitions, false-positive checks, and solution knowledge.

Source of truthReference onlyNo proof leakage
Key question
Which material can prove the organization’s present state, and which material may only guide interpretation?
Reasoning logic
Use uploaded sources to establish current-state claims. Use criteria, anti-patterns, tactics, taxonomies, and reference documents only as controlled analytical context.
Grounding basis
Taxonomy usage boundaries, criteria definitions, tactic database rules, and remote-reference-KB runtime status.
Reliability control
The Knowledge Base cannot supply a customer quote, justify a current-state score, or fill a missing document with best-practice assumptions.
Illustrative thinking sample“The KB defines what mature tagging looks like. It does not prove that the assessed organization has implemented it.”

2. Source registry, chunking & domain packetization

The complete source pack is converted into traceable chunks and routed into six bounded FinOps domain packets before parallel audit begins.

Context is deliberately bounded
Source IDsPage / row / chunk IDsA–F relevance routing35k target / 45k hard cap
ID
Open logic
Traceability layer

Source registry & provenance

Assigns stable source, page, chunk, row, and image identifiers to the assessment material.

Source IDChunk IDManifest
Key question
Can every later quote, downgrade, and finding be traced back to an exact source location?
Reasoning logic
Build a source registry that persists the location and type of each evidence unit before domain reasoning begins.
Grounding basis
Document boundaries, PDF page markers, table rows, character spans, image metadata, and packet manifests.
Reliability control
A plausible statement without a traceable registry location cannot function as strong evidence merely because it sounds domain-relevant.
Illustrative thinking sample“The evidence verifier should be able to point to src-003, page 007, chunk 2—not only to a generic document title.”
A–F
Open logic
Deterministic routing

Domain relevance classification

Scores each source chunk against domain-specific terms before the LLM receives its packet.

High / medium / lowKeyword reasonsPre-routing
Key question
Which chunks are materially relevant to Cost Visibility, Optimization, Governance, Architecture, Culture, or GenAI Cost?
Reasoning logic
Use deterministic domain signals to rank chunks and reduce irrelevant context before model interpretation.
Grounding basis
Domain-specific term sets, source names, gap terms, contradiction terms, and routing scores.
Reliability control
Each domain receives a tailored evidence universe instead of the full source dump, reducing cross-domain contamination and false pattern matching.
Illustrative thinking sample“Token-routing evidence belongs primarily in Domain F; a governance policy may also route to C, but it should not dominate right-sizing analysis in B.”
P
Open logic
Bounded evidence packet

Packet construction & weak coverage

Selects the highest-value chunks for each domain and marks thin coverage explicitly.

Bounded packetGap signalsWeak-coverage flag
Key question
What is the smallest sufficiently rich evidence set each domain auditor should receive?
Reasoning logic
Prioritize high and medium relevance, retain gap and contradiction signals, and stop near the packet target before context becomes noisy.
Grounding basis
Routing tiers, packet character budgets, included-chunk manifests, coverage notes, and image routing.
Reliability control
Thin packets are labeled as weak coverage. The system does not reinterpret packet silence as proof of maturity or proof that an anti-pattern is absent.
Illustrative thinking sample“A Domain E packet with only one generic culture reference is not enough to certify engineering cost accountability; broad-source fallback remains visibly marked.”
RT
Open logic
Execution control

Task-fit model routing & fallback

Maps each LLM task to a controlled primary profile and ordered fallback chain.

Stage IDsTask fitRun trace
Key question
Which reasoning capability is appropriate for pre-flight, audit, verification, rescan, synthesis, and fact-checking?
Reasoning logic
Use stronger or more independent model profiles where the task is consequential, and preserve ordered provider fallbacks rather than allowing free model choice.
Grounding basis
Central model registry, stage-specific profiles, normal vs. cheap-test mode, ordered fallbacks, and recorded actual model usage.
Reliability control
Models do not select their own role. Routing policy governs which model may perform each task and records what actually ran.
Illustrative thinking sample“A targeted rescan is routed differently from initial extraction because it must challenge a disputed finding rather than repeat the same first-pass behavior.”

3. Six parallel forensic domain audits

Each domain audits five maturity criteria and five corresponding anti-patterns. The two streams remain separate: capability evidence does not erase harmful patterns, and silence does not become maturity.

Parallel dual-stream reasoning
Shared audit contract: 0–3 evidence scale, exact source quotes, seven evidence categories, no benefit of the doubt, and explicit silence. Across six domains, the Engine evaluates 30 maturity criteria, 30 anti-patterns, and 180 evidence questions.
A
Open logic
Forensic domain

Cost Visibility & Allocation

Tests whether cloud spend is attributable, timely, role-visible, and connected to business unit economics.

TaggingShowbackUnit economics
Expanded sample model
Key questions
Can spend be attributed to owners, products, and cost centers? Are anomalies visible quickly? Do dashboards serve engineering, finance, and executives? Is cost translated into business units?
Reasoning logic
Evaluate the maturity stream against visibility anti-patterns such as tag sprawl, black-box spend, delayed reporting, siloed views, and vanity metrics.
Grounding basis
Tagging policies, allocation reports, dashboards, cost-center mappings, anomaly records, and unit-cost measures.
Reliability control
A written tagging policy scores differently from enforced tagging with coverage reporting and remediation evidence. A dashboard screenshot does not automatically prove accepted chargeback.
Critical note for accuracy
Visibility is often overstated because cost data exists somewhere. The audit asks whether the right people can use it for accountable decisions.
Illustrative thinking sample“A monthly finance export proves some reporting. It does not prove near-real-time anomaly response, engineering self-service, or reliable product-level allocation.”
B
Open logic
Forensic domain

Rate & Usage Optimization

Tests continuous commitment, right-sizing, waste, spot, and storage-lifecycle optimization.

CommitmentsRight-sizingWaste
Key question
Is optimization a continuous operating discipline or a periodic manual exercise?
Reasoning logic
Compare functioning optimization mechanisms with commitment avoidance, chronic over-provisioning, zombie resources, manual-only optimization, and discount blind spots.
Grounding basis
Commitment coverage, utilization reviews, automated recommendations, waste reports, shutdown schedules, and realized-savings tracking.
Reliability control
Recommendations in a tool do not prove that teams review, execute, and measure the actions. Claimed savings require a defensible baseline.
Illustrative thinking sample“A right-sizing report without ownership or completed actions is evidence of opportunity visibility, not evidence of embedded optimization.”
C
Open logic
Forensic domain

Governance & Policy

Tests policy, budgeting, operating model, procurement, and compliance-cost controls.

PolicyForecastingRACI
Key question
Do governance mechanisms influence real spending behavior, or do they exist mainly as documents and meetings?
Reasoning logic
Compare governed budgeting and accountability with shadow accounts, tolerated overruns, FinOps theater, vendor-lock-in blindness, and compliance treated as an afterthought.
Grounding basis
Policies, budgets, forecast variance, approval rules, RACI, governance cadence, procurement evidence, and compliance cost controls.
Reliability control
Policy evidence alone is categorized as policy. Higher maturity requires process, accountability, operational, or automation evidence showing the policy has teeth.
Illustrative thinking sample“A cloud policy can establish intent; exception logs, budget decisions, and enforcement mechanisms establish operating proof.”
D
Open logic
Forensic domain

Architecture & Engineering

Tests whether cost is embedded in design, IaC, scaling, cloud strategy, and platform choices.

Cost-aware designIaC guardrailsScaling
Key question
Does architecture prevent avoidable cost, or does FinOps mainly react after engineering decisions have already created spend?
Reasoning logic
Compare cost-aware architecture and automated guardrails with lift-and-shift, cost-blind design, unlimited scaling, single-cloud tunnel vision, and monolith tax.
Grounding basis
Architecture review templates, IaC policies, autoscaling limits, container/serverless practices, platform standards, and cost-performance decisions.
Reliability control
A target architecture or cloud strategy does not prove production implementation. Evidence must show decisions, controls, releases, or operational behavior.
Illustrative thinking sample“Autoscaling exists, but without upper bounds or cost alarms it may strengthen performance while preserving a scaling-without-limits anti-pattern.”
E
Open logic
Forensic domain

Culture & Organization

Tests ownership, executive sponsorship, engineering accountability, collaboration, and continuous improvement.

OwnershipCollaborationBehavior change
Key question
Who owns cloud economics, and does that ownership change daily engineering and financial behavior?
Reasoning logic
Compare an enabling FinOps operating culture with “cost is IT’s problem,” blame-based management, lip service, finance-engineering walls, and static maturity assumptions.
Grounding basis
Team charters, decision rights, review cadence, engineering KPIs, training, executive sponsorship, and cross-functional working practices.
Reliability control
Culture is not inferred from aspirational values. The audit looks for recurring mechanisms, accountability, incentives, and evidence of changed decisions.
Illustrative thinking sample“A FinOps community of practice is useful, but it does not prove engineering cost accountability unless teams carry measurable ownership.”
F
Open logic
Forensic domain

GenAI & AI Cost Management

Tests token visibility, AI allocation, routing efficiency, guardrails, forecasting, and value realization.

Token spendModel routingAI unit economics
Current sixth domain
Key questions
Can AI spend be attributed by use case and owner? Are premium models used selectively? Are prompt/context costs controlled? Are AI budgets, value hypotheses, and governance connected?
Reasoning logic
Compare functioning AI cost controls with invisible token spend, playground-to-production cost drift, premium-model overuse, unbounded context growth, and AI value theater.
Grounding basis
Model and token telemetry, use-case allocation, routing rules, prompt/context optimization, AI budgets, guardrails, and value-realization evidence.
Reliability control
Using an AI API does not prove AI FinOps maturity. Higher scores require attribution, governance, optimization, budgeting, and demonstrated value.
Critical note for accuracy
AI cost is unusually easy to hide in API usage and application budgets. Domain F prevents a traditional cloud-cost assessment from missing a rapidly growing cost surface.
Illustrative thinking sample“A premium model used everywhere may increase quality, but without routing policy and task-level unit economics it can indicate premium-model overuse rather than deliberate optimization.”

4. Independent evidence check, anti-pattern semantics & targeted rescan

The first audit is provisional. A separate verification layer checks whether forwarded scores and quotations are genuinely supported before any metric is calculated.

The audit checks itself
Provisional findingSupported / weak / unsupported / missingTargeted rescanDeterministic downgrade
V
Open logic
Independent verifier

Claim and score verification

Tests each forwarded criterion against the raw source packet and exact evidence location.

SupportedWeakUnsupported / missing
Key question
Does the source support both the finding and the strength of the assigned 0–3 score?
Reasoning logic
Verify the quote, surrounding context, packet coverage, and score calibration independently from the forensic auditor.
Grounding basis
Raw source chunks, exact IDs, original and verified counts, quote support, and coverage reasoning.
Reliability control
The verifier can lower the score even when the initial auditor produced fluent reasoning. Scoring authority remains conditional on source support.
Illustrative thinking sample“The source mentions monthly cost review, but the scanner scored continuous optimization at 3. Reclassify as weak and lower the verified count.”
Ø
Open logic
Absence semantics

Anti-pattern adjudication

Separates harmful findings, partial signals, tested absence, and unassessed silence.

Confirmed presentPartially presentTested / unknown absence
Expanded reliability model
Key questions
Is the harmful pattern present? Is the signal only partial? Did the source genuinely cover the risk area well enough to prove absence? Or is absence simply unknown?
Reasoning logic
Give negative findings their own semantics so a zero score cannot automatically become a positive maturity signal.
Grounding basis
Verifier status, original and verified counts, source coverage reason, disputed-signal adjudication, and weak-packet markers.
Reliability control
“Tested absent” is forbidden when the source is silent, irrelevant, weak, contradictory, or already contains a partial harmful signal.
Critical note for accuracy
This prevents the classic assessment error: interpreting missing evidence of a problem as evidence that the problem does not exist.
Illustrative thinking sample“No mention of commitment avoidance is not tested absence. A detailed commitment-management review showing rational coverage decisions can support tested absence.”
Open logic
Second opinion

Targeted rescan

Re-examines only weak, unsupported, or missing scored criteria instead of rerunning the whole domain blindly.

Disputed criteriaFocused promptHigher-value review
Key question
Can a narrower, stronger second pass locate evidence or correct an over-strong first-pass interpretation?
Reasoning logic
Feed the disputed criterion IDs and verifier feedback into a targeted audit task, then run the evidence check again.
Grounding basis
Original packet, verifier rationale, original count, disputed criteria list, and dedicated targeted-rescan routing.
Reliability control
Rescan is not a vote that automatically restores the first score. The second result must still pass independent evidence verification.
Illustrative thinking sample“Re-open only B2 and B4 because the verifier found related material but insufficient proof; do not regenerate all ten Domain B judgments.”
Open logic
Deterministic correction

Apply verified counts

Writes evidence-check outcomes back into the audit logs before Phase 2 sees them.

Original vs. verifiedAdjustment reasonNo optimism carryover
Key question
What score and absence status is the deterministic layer permitted to pass forward?
Reasoning logic
Replace unsupported or over-scored counts, preserve adjustment reasons and rescan status, and merge all domain verification results.
Grounding basis
Evidence-check items, anti-pattern semantics, adjustment records, and merged batch summaries.
Reliability control
Metrics are calculated from verified counts—not from the initial model’s preferred interpretation.
Illustrative thinking sample“Scanner score 3, verifier score 1, targeted rescan still weak: Phase 2 receives 1 and the report retains the downgrade trail.”

5. Deterministic metric firewall & confidence bracket

AI does not decide the headline maturity result. Arithmetic converts the verified audit into bounded metrics, classification, and permission for later synthesis.

No generative scoring
Core rule: maturity depth is reduced by confirmed anti-pattern burden, can receive only a small tested-clearance bonus, and is capped when verified evidence density is sparse.
Σ
Open logic
Metric firewall

Evidence-gated FinOps readiness

Calculates maturity, burden, clearance, coverage, integrity, density, and domain scores from verified audit data.

Deterministic math0–100Traceable inputs
Key question
What overall readiness does the verified evidence mathematically permit?
Reasoning logic
Normalize 30 maturity and 30 anti-pattern scores, subtract 50% of burden, add a small clearance bonus only when anti-pattern coverage is sufficient, and apply evidence caps.
Grounding basis
Verified counts, tested absences, evidence quotes, category footprints, delivered criteria, and domain totals.
Reliability control
No model can “round up” the final score because the organizational story sounds promising.
Illustrative thinking sample“Strong policy maturity with entrenched manual optimization and weak evidence density cannot become a high readiness score through narrative synthesis.”
C/W/R
Open logic
Classification

Crawl / Walk / Run

Translates the capped readiness score and burden into an evidence-aware maturity classification.

InsufficientCrawl / WalkRun
Key question
What maturity label is defensible after evidence density and anti-pattern burden are considered?
Reasoning logic
Below the evidence floor, classification is “Insufficient evidence.” Otherwise, readiness bands separate Crawl, Walk, Walk with significant friction, and Run.
Grounding basis
FinOps readiness, anti-pattern burden, and verified evidence density.
Reliability control
Low apparent burden cannot improve classification when anti-pattern coverage is weak; absence may simply be unknown.
Illustrative thinking sample“A 55 readiness score with burden above 50 becomes ‘Walk with significant friction,’ not a clean Walk.”
CAP
Open logic
Evidence cap

Evidence density & readiness ceiling

Caps optimistic readiness when too few criteria have verified source coverage.

<30 BLOCK floor<60 warning capCoverage matters
Key question
How much confidence can the source pack carry, independently of the apparent maturity score?
Reasoning logic
Below 30% density, readiness is capped by available evidence and the Quality Gate blocks action. Between 30% and 60%, readiness remains capped until stronger evidence is supplied.
Grounding basis
Criteria with verified source quotes, quote-backed gaps, findings, or meaningfully tested anti-pattern absence.
Reliability control
The Engine refuses to convert sparse documentation into false maturity certainty.
Illustrative thinking sample“A polished cloud strategy covering only eight of sixty evidence surfaces cannot justify a high enterprise FinOps readiness result.”
H/M/L
Open logic
Synthesis permission

Confidence bracket

Converts density, delivery integrity, and silent areas into HIGH, MEDIUM, or LOW synthesis behavior.

HIGH directiveMEDIUM cautiousLOW findings-only
Permission model
Key questions
Is there enough verified evidence and pipeline completeness to prescribe action? Should the Engine be directive, cautious, or stop at findings and validation needs?
Reasoning logic
HIGH requires at least 70% evidence density, 95% delivery integrity, and no more than 10 silent maturity areas. LOW is triggered below 30% density, below 70% integrity, or above 18 silent areas. Everything between is MEDIUM.
Grounding basis
Deterministic Phase 2 metrics only.
Reliability control
The synthesis model cannot choose a more ambitious mode. Evidence permissions are decided before the model call.
Critical note
Low-confidence runs still produce a useful evidence and validation report; they do not receive case studies, directive tactics, or a fabricated roadmap.
Illustrative thinking sample“The organization may have real gaps, but if the source pack cannot ground them deeply, the correct output is a validation plan—not confident implementation directives.”

6. Evidence summary, diagnosis & persona lenses

The Engine first establishes a facts-and-metrics summary, then interprets root causes, and only then translates the same evidence through three executive lenses.

Explanation after validation
ES
Open logic
Facts-first synthesis

Assessment evidence summary

Creates the non-prescriptive synopsis of classification, metrics, strengths, gaps, anti-patterns, and missing evidence.

Evidence onlyNo tactic IDsNo directives
Key question
What can the report say directly from verified evidence and locked Phase 2 metrics?
Reasoning logic
Build a factual assessment layer before causal interpretation or planning language is introduced.
Grounding basis
Crawl/Walk/Run classification, locked metrics, verified strengths, gaps, anti-patterns, and silent areas.
Reliability control
Case studies, tactics, invented outcomes, and implementation recommendations are excluded from this layer.
Illustrative thinking sample“The report may state that evidence density is 58% and anti-pattern coverage is 42%; it cannot translate those values into invented annual savings.”
D
Open logic
Interpretive layer

FinOps diagnosis

Explains the primary bottleneck, root causes, domain implications, and confidence without yet prescribing the roadmap.

Primary bottleneckRoot causesA–F diagnosis
Expanded sample model
Key questions
What evidenced mechanism best explains the current state? Which gaps are causal rather than merely correlated? Where does uncertainty remain?
Reasoning logic
Connect verified patterns across domains without changing the locked score or inventing organization facts.
Grounding basis
Evidence summary, category scores, anti-pattern findings, verified absences, source gaps, and confidence bracket.
Reliability control
Diagnosis may explain “why,” but it cannot prescribe “how” or create new numbers, owners, or behaviors not found upstream.
Critical note for accuracy
Cross-domain coherence is useful only when each link remains traceable. Narrative smoothness must not turn separate weak signals into a false causal certainty.
Illustrative thinking sample“Weak allocation, finance-engineering separation, and manual optimization may jointly indicate an accountability bottleneck—but the diagnosis must retain uncertainty when ownership evidence is sparse.”
FL
Open logic
Persona lens

FinOps Lead

Reads the evidence through operational maturity, optimization flow, tooling, and cross-team enablement.

OperationalMaturity gapsTeam enablement
Key question
Which verified operating mechanisms and gaps most affect the FinOps practice owner?
Reasoning logic
Translate the same factual diagnosis into FinOps terminology and operational priorities without changing the facts.
Grounding basis
Locked evidence summary and diagnosis.
Reliability control
The persona changes emphasis and vocabulary, not scores, source truth, or confidence.
Illustrative thinking sample“Emphasize ownership cadence, optimization backlog flow, and evidence gaps—not an invented tooling shopping list.”
CFO
Open logic
Persona lens

CFO / Finance Director

Reads the same diagnosis through financial control, risk exposure, budgeting, and investment confidence.

FinancialRiskInvestment logic
Key question
What does the verified evidence mean for financial predictability, accountability, and decision risk?
Reasoning logic
Use business and financial language while preserving the same evidence boundaries and avoiding unsupported ROI claims.
Grounding basis
Budgeting, allocation, unit economics, governance, financial-integration evidence, and locked metrics.
Reliability control
The CFO lens cannot invent annual savings, payback periods, or investment amounts absent from the source.
Illustrative thinking sample“Explain cost predictability and governance risk; do not claim a €2M savings opportunity unless a validated source and calculation support it.”
CTO
Open logic
Persona lens

Engineering Lead / CTO

Reads the evidence through architecture, engineering workflow, technical debt, and cost-performance trade-offs.

ArchitectureWorkflowCost-performance
Key question
Which evidenced engineering and architecture mechanisms are creating or constraining cloud cost efficiency?
Reasoning logic
Translate the same diagnosis into technical consequences while keeping maturity and anti-pattern facts unchanged.
Grounding basis
Architecture reviews, IaC, autoscaling, platforms, GenAI routing, and engineering-accountability evidence.
Reliability control
Technical specificity must come from the source or the verified tactic library; it cannot be improvised as if every cloud estate uses the same stack.
Illustrative thinking sample“Describe the evidenced absence of cost guardrails in IaC; do not assume Terraform, Kubernetes, or a specific cloud provider unless the source establishes it.”
Open logic
Adaptive routing

Synthesis escalation

Routes especially complex or high-friction diagnoses to a deeper synthesis stage using deterministic triggers or explicit deep mode.

Complexity triggersDeep modeRecorded reason
Key question
Is the organization’s evidence pattern complex enough to justify deeper synthesis effort?
Reasoning logic
Escalate when readiness is very low, burden is very high, gaps are numerous, or Crawl status contains many anti-patterns; a user can also explicitly request deep mode.
Grounding basis
Readiness, burden, gap count, anti-pattern count, maturity class, and user-controlled mode.
Reliability control
Deeper reasoning does not relax evidence rules. It changes analytical effort, not the source of truth.
Illustrative thinking sample“A low-readiness estate with widespread anti-patterns may justify deeper causal synthesis, but it still cannot bypass the LOW confidence findings-only boundary.”

7. Planning decision, roadmap synthesis & tactic permissioning

Recommendations are created only after diagnosis and only in the form permitted by the confidence bracket. The roadmap must remain linked to verified findings and approved tactic identifiers.

Strategy is permissioned
H/M/L
Open logic
Roadmap mode

Directive, cautious, or findings-only

Changes the shape of Phase 3 output according to evidence permission.

HIGH roadmapMEDIUM assumptionsLOW validation plan
Key question
What kind of action output is safe at the current confidence level?
Reasoning logic
HIGH can produce directive phases and tactics. MEDIUM produces cautious phases with confidence and assumptions. LOW produces evidence-backed findings, candidate themes, missing evidence, and a validation plan—without a directive roadmap.
Grounding basis
Precomputed confidence bracket and locked diagnosis.
Reliability control
The system does not force every assessment into a polished transformation roadmap.
Illustrative thinking sample“When evidence is LOW, ask for commitment coverage, allocation accuracy, and ownership records before prescribing implementation.”
GO?
Open logic
Actionability gate

Planning decision

States whether the roadmap is safe to use, conditionally usable, or blocked pending evidence.

GOCONDITIONAL_GONO_GO
Key question
Which actions are safe to execute now, and what evidence is required before the rest becomes actionable?
Reasoning logic
Separate diagnosis from prognosis: a valid diagnosis can coexist with a NO_GO planning decision when the evidence is not strong enough for implementation.
Grounding basis
Evidence summary, diagnosis confidence, roadmap mode, assumptions, and unresolved gaps.
Reliability control
Planning status is explicit. Readers are not left to infer actionability from the fluency of the report.
Illustrative thinking sample“Safe now: collect allocation and commitment evidence. Unsafe now: launch a chargeback redesign before ownership and allocation quality are validated.”
TAC
Open logic
Tactic permissioning

Verified tactic ID contract

Requires exact approved tactic IDs for prescriptive mechanisms and rejects invented or modified identifiers.

Exact IDsLookup tableRegenerate on invalid
Expanded control model
Key questions
Does the recommended mechanism exist in the verified tactics database? Does the locked finding match the tactic’s intended problem pattern? Is the identifier exact?
Reasoning logic
Place the complete tactic-ID lookup ahead of strategy generation, scan the resulting JSON for invalid IDs, and regenerate when the model invents, abbreviates, or modifies an identifier.
Grounding basis
Verified tactics database, tactic activity playbook, taxonomy registry, valid-ID set, and locked findings.
Reliability control
A plausible tactic name is not enough. IDs must exist exactly, and later tactic-grounding sanitation can replace or remove citations whose problem pattern is absent.
Critical note for accuracy
The tactic library is not evidence of need. It is a permissioned response library used only after the diagnosis establishes the relevant pattern.
Illustrative thinking sample“If a roadmap action cites a non-existent TAC-CUL shortcut or a tactic unrelated to the locked findings, regenerate or remove it.”
GND
Open logic
Finding-to-action control

Roadmap grounding sanitation

Removes unsupported actions and tactic citations after generation if they do not match both the action and the locked findings.

Finding corpusReplace / removeWarning trail
Key question
Does each action respond to a problem that is actually present in the verified diagnosis?
Reasoning logic
Compare action language and cited tactic rules against the locked finding corpus; replace a mismatched tactic or remove the unsupported action.
Grounding basis
Maturity gaps, anti-pattern findings, diagnosis, tactic-specific keyword rules, and roadmap action text.
Reliability control
Recommendation quality is checked after generation instead of assuming the prompt prevented every mismatch.
Illustrative thinking sample“A commitment-purchase tactic cannot remain in the roadmap if the source shows no stable workload pattern or commitment-management gap.”

8. Fact-check, sanitation, Quality Gate & audit trail

The complete report is treated as another artifact to verify. Unsupported claims are challenged, corrected, removed, or blocked before the final publication state is assigned.

Trust is earned after generation
FC
Open logic
Independent verification

Summary & roadmap fact-check

Reviews narrative claims and planning claims separately against sources, metrics, and approved tactic knowledge.

Separate checksBounded retriesHigh-reasoning escalation
Key question
Which generated claims are supported, misclassified, tactic-hygiene issues, materially unsupported, or unsafe to act on?
Reasoning logic
Check evidence summary and diagnosis separately from planning decisions and roadmap actions, then run bounded regeneration and a stronger retry when blockers survive.
Grounding basis
Source document, audit metrics, locked findings, tactic database, claim locations, and failure-type taxonomy.
Reliability control
A good score cannot hide an unsupported roadmap, and a strong roadmap does not excuse fabricated narrative facts.
Illustrative thinking sample“The diagnosis may be valid while one roadmap action remains unsupported. Check and repair them independently.”
SAN
Open logic
Active correction

Strategy sanitation

Removes, rewrites, or quarantines unsupported content while preserving an appendix trail.

RemoveRewriteQuarantine
Key question
Can the unsafe claim be corrected without changing the evidence contract, or must it disappear from the reader-facing report?
Reasoning logic
Operate on exact claims and source locations. Rewrite known metric misuse, remove unsafe roadmap actions, and quarantine unsupported fragments that can be located in the generated structure.
Grounding basis
Fact-check severity, failure type, claim text, source location, and deep strategy structure.
Reliability control
Sanitation cannot invent replacement evidence. It can only narrow, correct, or remove what the current evidence does not support.
Illustrative thinking sample“Rewrite ‘anti-pattern burden equals share of cloud spend’ into the correct interpretation: a validated severity index, not a spend allocation metric.”
QG
Open logic
Publication control

GO / WARN / BLOCK Quality Gate

Aggregates evidence density, traceability, verification, fact-check, silence, and sanitation into a visible final decision.

GOWARNBLOCK
Final authority
Key questions
Is the score valid? Is the strategy grounded? Are there material unsupported claims? Is the report safe to act on, usable with warnings, or blocked?
Reasoning logic
Block below the evidence floor or when material unsupported claims survive. Warn for partial evidence, weak anti-pattern coverage, many silent areas, sanitation, or incomplete fact-checking. Otherwise issue GO.
Grounding basis
Phase 1 and Phase 3 validators, evidence check, Phase 2 metrics, fact-check, sanitation record, and deterministic thresholds.
Reliability control
An explanatory model may clarify a WARN or BLOCK decision, but it cannot change the deterministic decision itself.
Critical note for use
A BLOCK does not necessarily invalidate every extracted fact or score. It means the final strategy is unsafe to act on until the listed evidence or grounding issues are resolved.
Illustrative thinking sample“Evidence density 28% triggers BLOCK even when the generated prose is excellent. The system chooses evidence sufficiency over presentation quality.”
TRACE
Open logic
Assurance layer

Run trace, diagnostics & drift corpus

Records the actual pipeline path and supports repeatability testing against bundled golden maturity profiles.

Model traceDiagnosticsGolden fixtures
Key question
Can the organization explain how this assessment ran and detect meaningful behavior drift over time?
Reasoning logic
Persist stage traces, model usage, warnings, packet status, evidence adjustments, and report diagnostics. Use synthetic Crawl, Walk, and Run fixtures as an MVP drift procedure.
Grounding basis
Run IDs, stage events, model results, source registry status, quality appendix, import/export diagnostics, and golden baseline files.
Reliability control
One successful run is not treated as proof of reproducibility. Drift tests expose score or narrative changes after model or criteria updates.
Illustrative thinking sample“After changing a model or rubric, rerun the known Crawl, Walk, and Run corpus and compare score and criterion behavior before production use.”

Evidence Summary

Facts, metrics, strengths, gaps, anti-patterns, and missing evidence.

FinOps Diagnosis

Primary bottleneck, root causes, and A–F domain interpretation.

Persona Views

FinOps Lead, CFO, and Engineering Lead lenses on the same evidence.

Roadmap or Findings Mode

Directive, cautious, or validation-first output according to confidence.

Quality Appendix

Evidence checks, sanitation, model trace, source gaps, and GO / WARN / BLOCK state.

Portfolio role: the Scanner can identify a potential company need; the FinOps Engine then performs a specialized evidence-gated diagnosis and creates only the action path the source material permits.
Interpretation boundaryThis visual shows the skeleton of FinOps thinking — not the full implementation
Open scope note
The presentation exposes the high-level reasoning architecture: how the Engine separates input safety, source provenance, dual-stream forensic audit, independent evidence checking, deterministic scoring, diagnosis, tactic permissioning, fact-checking, and publication control. The production system is built around this skeleton through lower-level contracts and operational mechanisms that are intentionally not shown here.

Reasoning implementation

  • Full prompts, role instructions, and criterion-level question wording
  • Exact orchestration, concurrency, retries, timeout handling, and state transitions
  • Structured-output schemas, repair procedures, and step-to-step contracts
  • Model-specific context shaping, reasoning effort, token limits, and fallback policy

Evidence infrastructure

  • Complete source-registry schemas, routing weights, packet manifests, and DLP rules
  • Evidence taxonomy implementation, provenance fields, quote validation, and page linking
  • Anti-pattern adjudication prompts, local support heuristics, and rescan prioritization
  • Remote Knowledge Base indexing, source boundaries, storage, and internal-result protection

Knowledge & decision systems

  • All 30 maturity criteria, 30 anti-pattern definitions, and 180 sub-questions
  • Complete tactic database, activity playbook, finding-keyword rules, and case material
  • Scoring formulas, thresholds, calibration choices, and golden-baseline content
  • Persona instructions, planning contracts, strategy guardrails, and taxonomy registry details

Operational assurance

  • Security controls, authentication, privacy sanitation, and report import/export safeguards
  • Exact fact-check taxonomies, regeneration appendices, and sanitation matching logic
  • Run-trace schemas, observability, internal model-result storage, and diagnostics
  • Regression scripts, drift procedures, benchmark evolution, and human validation practices
Design principle: begin with the intended decision and model the expert reasoning system around it. AI is used where probabilistic extraction, verification, diagnosis, and synthesis add value; deterministic controls govern source lineage, score calculation, confidence permissions, tactic validity, correction, and publication. The objective is to minimize free model trust—not to maximize autonomous model judgment.