AI Transformation Engine · Architecture Document

AI Transformation Assessment Engine — Evidence-Gated Scanner Architecture

This engine is not an AI tool checklist. It is a structured analytical pipeline for testing whether an organization can sense, decide, design, deliver, govern, learn, and scale AI-enabled value creation safely. AI readiness is organizational and architectural readiness, not tool readiness.

The engine evaluates actual operating evidence: decision rights, service-area ownership, brownfield integration, AI lifecycle control, semantic data context, value framing, embedded governance, safety controls, learning loops, and business architecture. Plans and slogans score low. Tool adoption alone does not prove readiness. Pilots become meaningful only when they are tied to value, safety, learning, and repeatable scaling.

The architecture is intentionally forensic. It uses local parsing, parse-quality checks, DLP scanning, five parallel audit batches, independent evidence verification, targeted rescans, deterministic metrics, confidence-gated synthesis, fact-checking, privacy-aware sanitation, and a GO/WARN/BLOCK quality gate.

AI readiness is the organization’s ability to turn AI into measurable value through adaptive operating models, reliable architecture, governed data, accountable service ownership, and safe learning loops.

The Forensic Analysis Pipeline

The pipeline separates extraction, verification, scoring, synthesis, and validation. The scanner extracts candidate evidence. The verifier checks whether the source material really supports it. The mathematics layer scores only verified data. The synthesizer receives locked metrics as boundaries. The fact-checker and sanitation layer then remove unsupported or privacy-unsafe claims before the report is shown.

Phase 0 — Pre-Flight Parsing and Safety
Local parsing · DLP scan · parse-quality metadata · relevance warnings
Client-side parsers extract source material before model analysis. PDF text is extracted with page markers, sparse or visual PDF pages can be carried as visual evidence, tables are normalized into summaries, HTML is sanitized, and JSON is normalized. A DLP pass plus deterministic redaction catches secrets and sensitive identifiers. Low AI-transformation keyword signal is a warning when material is parseable and safe, not a hard rejection.
source pack becomes bounded assessment input
Phase 1 — Five-Batch Forensic Audit
A-E batches · 50 criteria · 150 evidence questions
Five batches run across the AI Transformation domain model. Each batch evaluates 10 criteria: 5 maturity indicators and 5 opposing anti-patterns. Every criterion carries 3 evidence questions scored on a 0-3 evidence scale. Every captured quote is tagged with one evidence category so the report can distinguish policy evidence from operational, architecture, data, governance, learning, or value evidence.
Batch AAdaptive Operating Model
Batch BEnterprise AI Architecture & Platform Readiness
Batch CAI Strategy, Governance & Value Realization
Batch DData Foundations, Ownership & Accessibility
Batch EBusiness Capability & Service Architecture
The forensic auditor extracts evidence; it does not write strategy. This keeps scoring independent from the desirability of a future roadmap.
scanner output is provisional until independently checked
Phase 1.5 — Independent Evidence Check and Targeted Rescan
supported · weak · unsupported · missing · tested absent
Each completed A-E batch is checked against the raw source material before Phase 2 calculates metrics. The verifier classifies each forwarded claim as supported, weak, unsupported, or missing. Anti-pattern absence is separately labeled as confirmed present, partial finding, tested absent, or not assessed. When evidence is weak, unclear, unsupported, missing, or likely over-scored, the engine can independently trigger a targeted rescan and spend its rescan budget where a narrower second look matters most.
verified evidence crosses the deterministic firewall
Phase 2 — Deterministic Metric Firewall
No generative AI · score caps · evidence density · anti-pattern semantics
Phase 2 applies arithmetic to evidence-checked audit logs. It produces readiness, maturity depth, anti-pattern burden, anti-pattern clearance, anti-pattern coverage, delivery integrity, evidence density, domain scores, confirmed findings, verified absences, not-assessable anti-patterns, and silent areas. Sparse evidence caps the result. Missing evidence never becomes hidden optimism.
strategy receives locked metrics, not raw model confidence
Phase 3 — Confidence-Gated Synthesis
findings mode · roadmap mode · Kickstart · Building the System
Output depends on confidence. HIGH and MEDIUM confidence can produce a planning decision and a four-phase roadmap. LOW confidence is an autonomous stop condition: the engine returns findings mode only, with evidence-backed findings, missing evidence, validation actions, NO_GO, and an empty directive roadmap. This avoids spending time and trust on advice the evidence cannot carry.
draft strategy is independently challenged
Fact-Check, Sanitation, and Quality Gate
claim support · privacy readability · tactic validation · GO / WARN / BLOCK
The fact-check layer tests whether claims can be traced to source evidence or validated metrics. Unsupported strategy fragments are removed, rewritten, or quarantined. Privacy sanitation rewrites report-facing prose into fluent anonymized language. The quality gate aggregates delivery integrity, evidence density, grounding, fact-check status, tactic validity, and sanitation results into a visible GO/WARN/BLOCK decision.
Interactive Dashboard and Standalone HTML Export
persona switcher · source coverage gaps · appendix diagnostics
The final artifact includes the scorecard, A-E domain diagnosis, evidence check section, source coverage gaps, anti-pattern semantics, persona-specific summaries, roadmap or findings mode, quality appendix, and report import/export data. The report carries its own audit trail.

Why This Order Matters

Extraction is separated from recommendation

The forensic auditor cannot inflate scores because a future roadmap would sound attractive.

Verification happens before scoring

Unsupported claims are downgraded or rescanned before deterministic metrics can see them.

Mathematics caps optimism

Readiness is bounded by evidence density, anti-pattern coverage, and delivery integrity.

Reference KB is not customer proof

The Knowledge Base can shape rubric interpretation and tactics, but it cannot prove current state.

The synthesizer cannot rewrite diagnosis

Phase 3 receives locked scores, gaps, and confidence bracket as hard boundaries.

The engine can refuse to plan

LOW-confidence runs are routed to findings mode rather than polished but unsafe advice.

Agentic Verification: The Engine Decides What the Evidence Allows

The engine is agentic within guardrails. It does not simply pass text from one model call to the next. It evaluates intermediate results, compares them against the source, and makes bounded process decisions about what should happen next. Those decisions are not open-ended autonomy; they are evidence-gated controls inside the assessment workflow.

Targeted Rescansecond look
When a criterion is weak, unclear, unsupported, missing, or likely over-scored, the engine can ask for a narrower second pass focused on the disputed evidence instead of rerunning the whole assessment blindly.
Rescan Budgetprioritized thinking
If many items are weak, the engine prioritizes the most consequential gaps. It spends deeper model effort where the verification delta is highest and records what was skipped by budget.
Deterministic Downgradeproof before score
If the rescan still fails to support a forwarded finding, the score is lowered before the mathematical validation layer calculates maturity, burden, confidence, or readiness.
Anti-Pattern Semanticsabsence is tested
The engine distinguishes confirmed anti-patterns, partial findings, tested absence, and not-assessed silence. It does not turn missing evidence into a green result.
No-Roadmap Decisionconfidence gate
When confidence is LOW, the engine refuses to produce a directive roadmap. It returns findings mode, missing evidence, and a validation plan instead of polished but unsafe advice.
Execution Fast Pathstop unnecessary loops
When the evidence is too thin for planning, the engine skips expensive roadmap and fact-check loops that would only decorate uncertainty. Heavy thinking is reserved for evidence that can carry it.
Strategy Sanitationclaim hygiene
Unsupported strategy fragments are removed, rewritten, quarantined, or surfaced before display. The engine can preserve a report while refusing unsafe claims inside it.
Quality GateGO / WARN / BLOCK
The final decision is based on evidence density, delivery integrity, grounding, fact-check status, and sanitation results. The engine can block its own output when the assessment is not safe to act on.
The engine thinks by testing its own intermediate conclusions. It can ask for a narrower second look, refuse to plan, and surface uncertainty instead of producing polished but unsafe advice.

Three Roles. Three Constraints.

The engine behaves like a small assessment team. Each role has a bounded mandate, a different stage responsibility, and a different failure mode. The roles are not decorative prompt personas; they are control surfaces that keep parsing, scoring, synthesis, and validation separate.

01
Security Pre-Flight and Input Controller
Phase 0 · DLP, parsing, source readiness, relevance warnings
Role
DLP gatekeeper, parser supervisor, and relevance sentinel. It protects the pipeline from unsafe input while preserving usable low-signal material for evidence-gated assessment.
Behaviour
Conservative about safety, careful about relevance. Secrets, tokens, and sensitive identifiers are redacted before model analysis. Low AI-keyword density becomes a warning rather than a rejection when the material is safe and parseable.
Model
gemini-3.5-flash primary, with gemini-2.5-pro and gpt-5.5 fallback support. Local PDF.js extraction, CSV/TSV parsing, HTML/JSON handling, parse-quality checks, and deterministic redaction happen outside the model.
Decision
Hard blocks are reserved for unsafe, unreadable, unsupported, over-limit, or corrupt files. Generic IT/process material can pass into the scanner as adjacent context; the evidence pipeline later decides whether it proves AI Transformation readiness.

AI Transformation Signal Scan

[1] Core AI Transformation Vocabulary
AI operating model, AI governance, AI lifecycle,
agent, agentic workflow, autonomous decision,
human-in-the-loop, prompt, model gateway,
evaluation, groundedness, hallucination, RAG,
retrieval quality, red teaming, guardrail,
AI value hypothesis, Impact Statement,
Launch-and-Learn, service area ownership

[2] Architecture, Data, and Safety Signals
API boundary, brownfield integration, service
boundary, model registry, prompt registry,
rollback, observability, drift, lineage,
semantic context, data product, data access,
privacy, masking, purpose limitation, Agent
Behavioral Contract, Secure-by-Design

[3] Value-Creation and Operating Signals
service blueprint, value stream, capability map,
service catalog, decision rights, portfolio
learning, kill / continue / pivot, business case,
value realization, adoption, learning loop,
Sense & Respond, platform-as-product,
Kickstart, Building the System

Keyword signal estimates likely source richness; it does not score readiness. Generic operating-model documents can still be valuable when they prove decision rights, service ownership, value streams, or governance mechanisms that directly connect to AI Transformation criteria.

02
AI Transformation Forensic Auditor
Phase 1 · A-E batch scoring · maturity and anti-pattern extraction
Role
A forensic investigator who treats every uploaded document as an untrusted source. It extracts only what can be traced back to source text or selected visual evidence.
Per-Batch Behaviour
Receives 10 criteria definitions per batch: 5 maturity indicators and 5 anti-patterns. Each criterion has 3 evidence questions. It returns structured JSON with score, reasoning, evidence category, quote, source reference, silent-area notes, and source gaps.
Model · Concurrency
claude-sonnet-4-6 primary in the real routing. The fallback chain is gpt-5.5 audit profile, then gemini-2.5-pro. Five A-E batches run concurrently; evidence check runs independently after each batch completes.

Behavioural Rules

  • You are an AI Transformation Forensic Auditor. Your job is to extract explicit proof of AI-enabled value-creation readiness from uploaded source material.
  • Do not give the benefit of the doubt. Plans, slogans, aspirations, and future-tense claims score 1 at most.
  • Tool presence is not readiness. A chatbot, model gateway, platform, RAG index, or pilot must show operating proof before it can support higher scores.
  • General IT/process maturity is adjacent context unless it directly proves AI operating rhythm, AI architecture, data readiness, governance, safety, value realization, or service-area AI ownership.
  • Silence is data. If the source is silent on a criterion, score it 0 and record the missing evidence instead of balancing the audit.
  • Every captured evidence quote must be tagged as Policy, Process, Operational, Automation, Accountability, Architecture, Data, Governance, Learning, or Value.
  • Anti-pattern absence is not automatically positive. It must be classified as confirmed present, partial finding, tested absent, or not assessed.
  • Scanner output is provisional. Evidence-check can downgrade scores and request targeted rescans when raw material does not support the forwarded finding.
03
Strategic Synthesizer and Roadmap Planner
Phase 3 · persona summaries · findings mode · roadmap · quality narrative
Role
A transformation architect constrained by confidence. At high or medium confidence it can structure a roadmap; at low confidence it produces findings, missing evidence, and validation needs instead of directive advice.
What It Receives
Locked Phase 2 metrics, evidence-checked findings, source coverage gaps, anti-pattern semantics, verified absences, the AI Transformation tactic database, and persona lenses for Transformation Lead, CIO/CTO or CDAO, and Business / Service Area Owner.
What It Produces
Evidence-only persona summaries, evidence summary, A-E domain diagnosis, planning decision, visual scorecard, findings mode when needed, and a four-phase roadmap split between Kickstart and Building the System when confidence allows.
Model
claude-sonnet-4-6 writes evidence summaries and diagnosis. Escalation and roadmap planning use claude-opus-4-7 first, with gpt-5.5 and gemini-3.1-pro-preview as fallbacks.

Behavioural Rules

  • Source of truth: every diagnostic claim must trace to Phase 2 metrics, evidence-checked findings, or uploaded source material.
  • Do not invent owners, scores, numbers, readiness level, business value, adoption, controls, or maturity.
  • LOW confidence means findings mode: no directive roadmap, no tactic prescriptions, planning decision NO_GO, and explicit missing-evidence guidance.
  • Use the Knowledge Base only as rubric, false-positive, risk-framing, and tactic context. It must never become current-state proof, evidence quote, or visible source reference.
  • Write in fluent anonymized prose from the start. Do not mention company names, organization names, Knowledge Base document names, source labels, or placeholder redaction tokens.
  • Roadmap actions must follow Kickstart then Building the System: diagnose, select domain, run vertical slice, validate value/safety, then integrate architecture, platform, data, teams, rhythm, governance, and learning.
  • Every prescribed tactic action cites a valid tactic ID only when the locked finding matches the tactic pattern. Generic evidence-gathering actions may omit tactic IDs.
  • Persona summaries must agree on facts; they differ only in lens, vocabulary, and emphasis.
The report has three executive lenses, but the facts must agree. The Transformation Lead needs operating-model truth. The CIO/CTO or CDAO needs architecture, platform, data, and governance implications. The Business / Service Area Owner needs impact, adoption, ownership, and value-stream consequences.

Two Streams. One Diagnosis.

Every domain has two opposing streams. Stream A measures maturity: evidence of capability, ownership, operating proof, architecture, governance, data quality, learning, and value realization. Stream B measures anti-patterns: evidence of friction, fragility, decision fog, brittle integration, context-poor data, unsafe autonomy, or tool adoption without value proof.

High Maturity + Low Anti-Patterns

Genuinely adaptive only when anti-pattern coverage is high and absence has been meaningfully tested.

High Maturity + High Anti-Patterns

Structured but inconsistent. The organization has mechanisms, but friction blocks scale or trust.

Low Maturity + High Anti-Patterns

Transformation work is needed before AI-enabled value creation can scale safely.

Low Maturity + Low Anti-Patterns

Usually insufficient evidence unless the absence of anti-patterns was directly tested by source material.

Confirmed findingThe anti-pattern is directly supported by source evidence.
Partial findingThe source shows credible symptoms but not enough proof for full severity.
Tested absentThe source meaningfully covers the risk area and shows the blocker is not present.
Not assessedThe source does not test the issue. Absence is unknown, not good.

Anti-pattern absence is positive only when verified. A zero anti-pattern score can mean tested absent, which is useful evidence of clearance, or not assessed, where the source simply did not cover the issue. Low burden with low coverage means unknown, not excellent.

The 50-Criteria AI Transformation Knowledge Base

Five domains. Each domain holds 5 maturity indicators paired against 5 anti-patterns. Each criterion carries 3 evidence questions. The result is a 50-criterion, 150-question scanner for whether AI can become safe, measurable, service-owned value creation.

Embedded AI economics lens. GenAI and AI cost management is not a separate Domain F in this engine. It is treated as supporting readiness evidence inside the existing A-E model: B3/B5 cover observability, routing, quota, caching, and platform cost controls; C2/C5 cover unit economics, value-vs-cost evidence, budgeting, forecasting, and investment decisions; A/D/E add lighter checks for demand cost profile, hidden review cost, retrieval cost, and cost-to-serve traceability. AI spend dashboards alone do not prove AI readiness.
A

Adaptive Operating Model

Can the organization route, decide, fund, learn, and scale AI work through an adaptable operating rhythm?

Maturity Indicators

A1
AI Operating Rhythm & Decision Rights
Strategy, demand intake, delivery, learning, and governance are connected through clear cadence and ownership.
  1. Is there a documented AI operating rhythm?
  2. Are decision rights assigned across business, IT, data, architecture, security, legal, and service areas?
  3. Are decisions traceable through cadence, approval records, backlog movement, or escalation logs?
A2
Adaptive AI Demand Routing
AI demand is classified and routed according to uncertainty, risk, value profile, cost profile, capacity, competence, and service ownership.
  1. Are demands classified as predictable, standardized, iterative, exploratory, or transformative?
  2. Are different delivery paths used for different demand types?
  3. Is routing connected to capacity, competence, service ownership, value, cost profile, and risk level?
A3
Evidence-Based Launch-and-Learn Model
Bounded vertical slices validate value and safety before scaling.
  1. Does transformation start from a safe experiment or vertical slice?
  2. Are pilot learnings used to refine operating model, architecture, data, and governance?
  3. Is there a transition from Kickstart validation to Building-the-System scaling?
A4
Shared Learning Architecture
AI learning spreads through communities, chapters, guilds, playbooks, incidents, and reusable practices.
  1. Are competence structures used to spread AI learning?
  2. Are pilots, failures, red-team results, and incidents converted into practices?
  3. Do business, AI, data, platform, and service teams learn together?
A5
Human-AI Work Redesign
AI changes work design, roles, handoffs, review points, review cost, rework cost, and human-in-the-loop responsibilities.
  1. Are initiatives designed to reallocate human capacity toward higher-value work?
  2. Are roles, handoffs, exception paths, and review points redesigned?
  3. Are employee experience, cognitive load, hidden review or rework cost, quality, and customer impact measured?

Anti-Patterns

AP-A1
AI Decision Fog & Governance Fat
Decisions stall or repeat because ownership, boundaries, and escalation logic are unclear.
  1. Are AI initiatives delayed because nobody clearly owns decisions?
  2. Do separate forums create repeated approvals instead of flow?
  3. Are ordinary design decisions escalated because boundaries are unclear?
AP-A2
One-Size-Fits-All AI Delivery
Standardized automation, exploratory GenAI, and transformational service change are forced through the same project model without value, risk, or cost-profile routing.
  1. Are all AI initiatives forced through one delivery model?
  2. Are exploratory cases managed with rigid plans that assume known requirements?
  3. Are predictable automation and uncertain agentic work mixed without value, risk, or cost-profile routing logic?
AP-A3
Pilot Purgatory
Pilots remain detached from production, ownership, safety, reusable learning, and value proof.
  1. Do pilots remain disconnected from production, ownership, or outcomes?
  2. Are pilots called successful without integration, safety, adoption, or value proof?
  3. Does each use case start from scratch instead of reusing patterns?
AP-A4
Fragmented AI Knowledge & Hero Culture
AI know-how sits with isolated experts and mistakes repeat because learning is not institutionalized.
  1. Is knowledge concentrated in a few experts or isolated teams?
  2. Are mistakes repeated because learning is not institutionalized?
  3. Do teams depend on individual heroics instead of shared practices?
AP-A5
Digital Taylorism & Workslop
AI makes fragmented tasks faster while increasing hidden review cost, rework, checking, coordination, and low-quality output.
  1. Is AI used mainly to speed fragmented tasks?
  2. Does AI create extra checking, rework, hidden review cost, hidden coordination, or workslop?
  3. Are people measured on activity rather than impact, learning, quality, and customer value?
B

Enterprise AI Architecture & Platform Readiness

Can AI be integrated, deployed, observed, secured, and reused through reliable technical and platform foundations?

Maturity Indicators

B1
Service-Boundary-Based AI Integration
AI systems integrate through stable interfaces rather than brittle point connections.
  1. Are integrations built through APIs, service boundaries, events, or governed connectors?
  2. Are legacy systems wrapped before agents depend on them?
  3. Are ownership, versioning, dependencies, and failure modes documented?
B2
AI Lifecycle & Release Control
Models, prompts, agents, tools, datasets, configurations, and evaluations are versioned and releasable.
  1. Are models, prompts, agents, tools, datasets, configurations, and evaluations versioned?
  2. Is there a build-test-deploy-monitor-rollback lifecycle?
  3. Can production behavior be traced to model, prompt, data, tool, and release version?
B3
AI Observability & Evaluation System
Production AI is monitored for quality, safety, token/model spend, latency, routing performance, retrieval cost, cost-per-output, groundedness, drift, and business impact.
  1. Are quality, safety, token/model spend, latency, routing performance, retrieval cost, cost-per-output, groundedness, drift, and business impact monitored?
  2. Are offline and online evaluation methods defined?
  3. Do alerts trigger when performance, data freshness, value, cost, or safety thresholds degrade?
B4
Secure-by-Design AI Trust & Safety Layer
Guardrails, access controls, red teaming, escalation, and behavioral boundaries are built into service design.
  1. Are prompt injection, leakage, toxic content, unsafe tool use, and adversarial scenarios tested?
  2. Are guardrails, filters, access controls, and escalation paths embedded?
  3. Is there an Agent Behavioral Contract defining forbidden behaviors?
B5
AI Platform as Product
Shared AI capabilities include model gateways, routing, caching, quotas, budget alerts, data patterns, templates, and cost-aware developer experience.
  1. Are shared AI services, model gateways, routing, caching, quota controls, budget alerts, data patterns, and templates offered as platform products?
  2. Are platform teams accountable to service-area teams as internal customers?
  3. Are adoption, flow, reliability, cost-aware developer experience, and value enablement measured?

Anti-Patterns

AP-B1
Legacy Labyrinth & Brittle AI Connectivity
AI depends on manual exports, copied files, scraping, fragile connectors, or unclear service boundaries.
  1. Do use cases depend on manual exports, screen scraping, or fragile connectors?
  2. Does changing one system break another because boundaries are unclear?
  3. Are integrations created project-by-project with no reusable architecture?
AP-B2
Notebook-to-Production / Prompt-to-Production Chaos
AI artifacts reach production without reproducible release control.
  1. Are models, prompts, or agents moved into production without controls?
  2. Are prompt changes made manually without review or rollback?
  3. Is it impossible to reproduce why an AI system behaved a certain way?
AP-B3
Black-Box AI Operations
Failures, invisible token spend, unmonitored model usage, cost surprises, and unsafe behavior are discovered through complaints or financial surprises.
  1. Are failures discovered mainly through complaints or business damage?
  2. Are there no baseline metrics for accuracy, token/model spend, cost-per-output, value, safety, or reliability?
  3. Are drift, retrieval degradation, hallucinations, cost anomalies, or unsafe outputs not systematically detected?
AP-B4
Safety Theater & Unbounded Autonomy
Guardrails exist on paper while agents lack tested behavioral limits and escalation boundaries.
  1. Are guardrails trusted without test evidence?
  2. Can agents access tools, systems, or data without clear boundaries?
  3. Are safety reviews late gates instead of built into service design?
AP-B5
Tool Fragmentation & Hidden Factory
Teams recreate ingestion, model access, embeddings, vector stores, retrieval, hosting, governance, evaluation, and monitoring stacks separately.
  1. Does every project build its own model, embedding, vector-store, retrieval, and monitoring stack?
  2. Are teams buying disconnected tools without standards, cost controls, budget alerts, or observability?
  3. Does fragmentation create duplicate AI spend, maintenance debt, and reduced innovation capacity?
C

AI Strategy, Governance & Value Realization

Is AI connected to purpose, value, governance, safety, investment logic, and measurable business outcomes?

Maturity Indicators

C1
Purpose-Driven AI Strategy
AI ambition is tied to strategy, customers, values, business model, and boundaries.
  1. Does the organization define why AI matters to strategy, customers, values, and business model?
  2. Are ambition areas connected to service areas, value streams, or strategic domains?
  3. Are where-not-to-use boundaries explicit?
C2
Impact-Driven AI Value Framing
Initiatives start from Impact Statements, baselines, value hypotheses, unit economics, and post-release measurement.
  1. Are initiatives started from explicit Impact Statements rather than tool ideas?
  2. Are outcomes defined through customer value, profitability, quality, flow, experience, or cost per workflow, proposal, ticket, conversation, document, assessment, or decision?
  3. Are value hypotheses tested and updated after pilots or releases?
C3
Embedded AI Governance Model
Governance enables safe value creation inside service, platform, security, data, and business ownership.
  1. Are governance responsibilities embedded into service areas, platform, security, data, and business ownership?
  2. Does governance enable safe value creation rather than only blocking risk?
  3. Are exceptions, approvals, risks, and controls auditable?
C4
Risk-Based Responsible AI Control
Use cases are classified by risk, autonomy, user impact, data sensitivity, and regulatory exposure.
  1. Are use cases classified by risk, autonomy, user impact, data sensitivity, and regulation?
  2. Are oversight, escalation, disclosure, logging, and accountability tied to risk?
  3. Are high-risk or autonomous scenarios blocked unless minimum controls are proven?
C5
Evidence-Based AI Investment Portfolio
Portfolio decisions use impact, risk, readiness, budgeting, forecasting, spend guardrails, value-vs-cost, and learning evidence.
  1. Are investments managed through impact, risk, readiness, budgeting, forecasting, spend guardrails, cost, and learning logic?
  2. Are kill, continue, scale, and pivot decisions based on value-vs-cost evidence?
  3. Does pilot learning reprioritize the portfolio?

Anti-Patterns

AP-C1
AI Slogan Strategy
AI is a generic transformation theme without choices, boundaries, or value logic.
  1. Is AI described generically without strategic choices?
  2. Is everything labeled an AI priority without prioritization logic?
  3. Are tools adopted before value is defined?
AP-C2
Use-Case Chasing & Vanity Benefits
AI Value Theater: initiatives are selected for fashion, visibility, tool adoption, token/model spend, or cost claims rather than measured impact.
  1. Are initiatives selected because they are visible or technically interesting?
  2. Are savings, ROI, TCO, token/model spend, or efficiency claims made without baseline and post-measurement?
  3. Are spend, usage, or successful pilots celebrated without enterprise-level impact?
AP-C3
Rigid Gatekeeping Governance
Governance slows flow without improving safety, causing bypass behavior or shadow AI.
  1. Does governance slow flow without improving safety?
  2. Are teams unclear how to proceed safely, causing bypass or shadow AI?
  3. Are governance decisions disconnected from impact, architecture, and operations?
AP-C4
Unclassified Risk & Ambiguous Accountability
AI is launched without risk classification, accountable owner, override, appeal, fallback, or escalation.
  1. Are use cases launched without risk classification?
  2. Is accountability unclear when AI-assisted decisions cause harm or failure?
  3. Are override, appeal, fallback, or escalation mechanisms missing?
AP-C5
AI Investment Drift
The backlog, platform usage, pilot estate, and AI spend grow without capacity, readiness, value proof, guardrails, or willingness to stop weak work.
  1. Does the backlog, platform usage, or AI spend grow without capacity, readiness, or value proof?
  2. Are weak initiatives kept alive by sunk cost, politics, or enthusiasm?
  3. Are investments focused on platform build-out, premium model usage, or tool expansion without near-term capability validation?
D

Data Foundations, Ownership & Accessibility

Is data usable as contextual, governed, service-owned fuel for AI decisions and workflows?

Maturity Indicators

D1
Domain-Owned AI Data Products
AI-critical datasets have owners accountable for quality, meaning, lifecycle, and usability.
  1. Are critical AI datasets owned by the service area or domain that understands context?
  2. Are owners accountable for quality, meaning, lifecycle, and usability?
  3. Are data products published with context, definitions, access rules, and usage expectations?
D2
Context-Rich AI-Ready Data
Data carries semantic meaning, source context, process linkage, and decision relevance.
  1. Does data include semantic context, source meaning, process linkage, and operating conditions?
  2. Can AI systems understand what values mean in service context?
  3. Are critical data elements mapped to decisions, triggers, workflows, or outcomes?
D3
Data Quality & Lineage Control
Training, retrieval, inference, and decision data are traceable, quality-checked, and freshness-monitored.
  1. Are critical datasets cataloged, versioned, quality-checked, and traceable?
  2. Are freshness, completeness, correctness, and drift monitored?
  3. Can training, retrieval, inference, and decision data be traced after the fact?
D4
Governed AI Data Access
Classification, purpose limitation, access rights, privacy, masking, minimization, and retention are designed into data flows.
  1. Are classification, purpose limitation, access rights, and privacy controls documented?
  2. Is access role-based, auditable, and aligned with ownership?
  3. Are masking, anonymization, minimization, retention, and boundary controls used where required?
D5
Reusable AI Data Access Patterns
Retrieval, feature, search, embedding, vector-store, batch, real-time, and event patterns are governed, reusable, and observable for context growth and retrieval cost.
  1. Are structured, unstructured, real-time, batch, feature, and retrieval patterns defined?
  2. Are RAG, embedding, search, feature-store, vector-store, and event patterns governed and reusable?
  3. Is retrieval quality, coverage, freshness, grounding, context growth, duplication, and retrieval cost measured?

Anti-Patterns

AP-D1
Ownerless Data & Context Loss
Data ownership is centralized, undefined, or detached from domain meaning.
  1. Is data ownership centralized in IT or undefined, causing context loss?
  2. Are AI teams forced to interpret datasets without domain meaning?
  3. Do data issues stall AI because nobody owns correction or interpretation?
AP-D2
Raw Data Without Meaning
AI receives available signals without the “why this data matters” layer.
  1. Is data collected because available rather than because it drives decisions?
  2. Do models receive raw signals without business, process, or operational context?
  3. Are outputs unreliable because the model lacks meaning and decision context?
AP-D3
Data Swamp & Broken Lineage
AI uses data whose origin, version, freshness, or quality cannot be proven.
  1. Are outputs based on data whose origin, version, or quality cannot be proven?
  2. Are stale copies, manual files, or undocumented transformations used?
  3. Are quality failures found only after model failure or escalation?
AP-D4
Data Leakage & Over-Broad Access
Sensitive data can be accessed or copied without purpose, logging, approval, or lifecycle control.
  1. Can tools access sensitive data without purpose, logging, or approval?
  2. Are copies moved into experimentation without lifecycle control?
  3. Are privacy and security controls added late?
AP-D5
Ad Hoc Retrieval & Duplicate Knowledge Stores
Teams build unowned indexes, embeddings, vector stores, and knowledge bases that become stale, inconsistent, duplicated, or costly.
  1. Does each team create separate embeddings, vector stores, or knowledge bases without ownership or cost visibility?
  2. Do RAG systems use stale, incomplete, or unverified sources?
  3. Are retrieval failures, unbounded context growth, and duplicated retrieval costs treated only as model failures?
E

Business Capability & Service Architecture

Are AI opportunities anchored in customer needs, service blueprints, value streams, business capabilities, and end-to-end ownership?

Maturity Indicators

E1
AI-Anchored Service & Capability Architecture
AI opportunities link to service catalogs, capability maps, customer paths, or value streams.
  1. Is there a current service catalog, capability map, or equivalent model?
  2. Are opportunities linked to services, capabilities, customer paths, or value streams?
  3. Are dependencies between capabilities, platforms, data, and outcomes visible?
E2
AI-Ready Service Blueprinting
Services are mapped end-to-end before AI interventions are placed.
  1. Are target services mapped across customer journey, frontstage, backstage, systems, data, and handoffs?
  2. Are bottlenecks, decisions, failure modes, and automation candidates visible?
  3. Are interventions placed where they improve the whole value stream?
E3
Traceable AI Solution Structure
Features trace from business need, outcome, and cost-to-serve to capability, component, data, platform, and work package.
  1. Are features decomposed into need, capability, component, data, platform, and work packages?
  2. Can each use case be traced to customer need, outcome, cost-to-serve where evidence exists, and service architecture?
  3. Are requirements, acceptance criteria, risks, and dependencies explicit before build?
E4
Service Area Ownership of AI Value Creation
Integrated service-area teams own business, application, data, and AI outcomes as living products.
  1. Are integrated teams accountable for business, application, data, and AI outcomes?
  2. Are platform teams positioned as enablers for service-area teams?
  3. Are AI solutions owned as living products rather than temporary projects?
E5
Phased AI Scaling Through Service Areas
Scaling is sequenced by service-area readiness, architecture, data, risk, cost-to-serve awareness, and learning evidence.
  1. Is scaling sequenced by readiness, architecture, data, risk, and cost-to-serve readiness where evidence exists?
  2. Are pilot patterns converted into playbooks, templates, platform capabilities, and guardrails?
  3. Is continuous improvement built into rollout through feedback and reassessment?

Anti-Patterns

AP-E1
AI Ideas Detached from Business Architecture
Use cases are not tied to service catalog, capability map, owner, dependencies, or value stream.
  1. Are use cases listed without connection to service catalog or capability map?
  2. Is it unclear which service area owns the change?
  3. Are dependencies discovered only during implementation?
AP-E2
Spot Optimization & Silo Automation
AI speeds one task while increasing downstream rework, coordination, or customer friction.
  1. Is AI applied to isolated tasks without surrounding flow?
  2. Does automation speed one step while increasing downstream work?
  3. Are customer experience and end-to-end flow ignored?
AP-E3
Untraceable AI Build Logic
Teams build components without traceability to need, capability, data, platform, outcome, or cost-to-serve where evidence exists.
  1. Are teams building AI components without linkage to capability or need?
  2. Are requirements vague, unstable, or tool-led?
  3. Is traceability between feature, data source, platform dependency, value outcome, and cost-to-serve missing?
AP-E4
Disconnected AI Project Teams
Temporary AI teams remain detached from service areas that must own outcomes.
  1. Are AI teams separate from the service areas that need to own outcomes?
  2. Are business, data science, application, and platform work sequential handoffs?
  3. Does ownership disappear after pilot completion or vendor delivery?
AP-E5
Big-Bang AI Transformation
AI scales before readiness, architecture, data ownership, cost-to-serve readiness, operating model, and safety evidence are proven.
  1. Is rollout scaled before architecture, data ownership, cost-to-serve readiness, and operating model are ready?
  2. Are pilot learnings not converted into reusable patterns?
  3. Does every service area repeat mistakes independently?

0-3 Evidence Scale and AI Maturity Ladder

The scanner asks what the submitted material proves. A statement that “AI is strategic” is not enough. A pilot is not enough. A platform purchase is not enough. Higher scores require operating proof, measured feedback loops, tested controls, ownership records, or traceable production behavior.

0
Insufficient evidence

No evidence found, or the source is not assessable for that criterion.

1
Emerging

Intent, plans, pilots, future-tense claims, or isolated examples without operating proof.

2
Structured

Functioning mechanism, ownership, process, control, or delivery practice is described.

3
Adaptive / value-creating

Mechanisms, measurement, enforcement, learning loops, and repeatable production behavior are visible.

Evidence Check Statuses

SupportedThe quote and surrounding context directly support the forwarded score or finding.
WeakThe source is relevant but too indirect, generic, partial, or ambiguous for the score claimed.
UnsupportedThe scanner overreached, inferred beyond the source, or mapped the evidence to the wrong criterion.
MissingThe criterion may matter, but the submitted material does not provide assessable evidence.

Readiness Classification

Insufficient evidenceblocked classification
The source material does not support a directive maturity conclusion. The report returns findings mode and source gaps.
Emergingearly capability
The organization shows intent, early pilots, or isolated mechanisms, but lacks enough operating proof to show repeatable AI value creation.
Structureddefined mechanisms
Core structures exist: ownership, governance, platform, data practices, and service framing are visible but not yet fluid or scaled.
Scaling with frictioncapability plus blockers
Meaningful AI transformation capability is visible, but anti-pattern burden or uneven service-area absorption blocks smooth scaling.
Adaptive / Value-creatinghigh evidence, low tested burden
AI-enabled value creation is service-owned, architecturally integrated, data-grounded, governed, measured, and continuously improved.

Evidence Categories

The score asks how strong the evidence is. The category tag asks what kind of evidence it is. Every evidence quote is tagged with exactly one category: Policy, Process, Operational, Automation, Accountability, Architecture, Data, Governance, Learning, or Value.

This matters because the assessed organization can have extensive policy evidence and almost no operating evidence. Another can have platform automation but no service ownership. AI cost signals are mapped into existing categories rather than a separate finance category: spend dashboards are Operational, routing and quotas are Automation, budget guardrails are Governance, owners are Accountability, and unit economics are Value.

Source Rules and Knowledge Base Guardrails

The engine treats the uploaded source pack as untrusted customer evidence and the remote Knowledge Base as internal methodology context. Those two sources are never allowed to collapse into one another.

{
  "source_evidence": "Only uploaded assessment material can prove current state.",
  "reference_kb": "Rubric context, validation questions, false-positive checks, risk framing, and tactics only.",
  "forbidden_kb_uses": [
    "customer_current_state_claim",
    "audited_finding",
    "evidence_quote",
    "score_justification_without_uploaded_source",
    "visible_document_name",
    "visible_company_or_organization_name"
  ],
  "general_process_boundary": "Generic IT/process maturity is adjacent context unless it directly proves an AI Transformation criterion.",
  "privacy_output_rule": "Write fluent anonymized prose from the start. Do not emit company names or placeholder tokens."
}
The Knowledge Base is a sparring partner for interpretation and tactics. It is not witness testimony about the assessed organization.

What to Submit

The assessment accepts any combination of documents. More document types provide richer signal across all five domains. Absence of entire document categories is itself a maturity signal. Do not manufacture missing documents for the assessment; absence is part of the diagnosis.

Tier 1 · Direct AI Transformation EvidenceAI operating model, AI governance standards, service-area AI ownership model, model/prompt lifecycle records, AI architecture diagrams, data product ownership, semantic data models, AI observability/evaluation records, red-team results, Agent Behavioral Contracts, pilot retrospectives with value and safety evidence.
Tier 2 · Strong Supporting EvidenceDecision-rights maps, service blueprints, value-stream maps, capability maps, portfolio reviews, risk registers, platform operating model, security and data access controls, impact dashboards, release records, learning loops, and operating cadence artifacts.
Tier 3 · Adjacent ContextGeneral project delivery, process governance, capacity management, sales/service operating models, role descriptions, scorecards, and business planning. These can support interpretation but do not prove AI readiness unless they directly connect to AI-specific operating, architecture, data, safety, value, or service-area evidence.
Local Parsingnot model parsing
PDF, HTML, CSV/TSV, JSON, and image inputs are parsed or encoded locally before analysis. Models analyze extracted material; they are not the primary parsers.
Visual Evidenceselected pages
Sparse-text or dashboard-like PDF pages can be included as images so visible diagrams, tables, operating models, service maps, and evidence screenshots are not lost.
Reference Knowledge Baserubric and tactics context
The remote Knowledge Base can provide methodology context and approved tactics. It is never treated as proof of the assessed organization’s current state.

The Firewall Between AI and Roadmap

Phase 2 contains no generative AI. It applies deterministic arithmetic to evidence-checked audit logs. Unsupported scanner output has already been rescanned or downgraded before this layer runs.

Evidence-Gated Readiness0-100, capped
readiness = maturity_depth - anti_pattern_penalty + tested_clearance_bonus then capped by evidence density and delivery integrity
The headline score starts from validated maturity, subtracts confirmed anti-pattern burden, rewards only tested clearance, and caps optimistic scores when source evidence is sparse.
Maturity Depth0-100
maturity_depth = (sum of 25 maturity scores / 75) * 100
Average maturity across all five domains. Captures partial progress that a simple “fully mature criteria” count would miss.
Anti-Pattern Burden0-100
burden = (sum of confirmed or partial anti-pattern scores / 75) * 100
Average severity across harmful evidence. Higher means more friction blocking safe AI-enabled value creation.
Anti-Pattern Clearance0-100
clearance = (tested absent anti-pattern criteria / 25) * 100
How much of the anti-pattern surface was meaningfully tested and found absent. This is the only way absence becomes positive evidence.
Anti-Pattern Coverage0-100
coverage = (tested_absent + confirmed_present + partial_finding) / 25 * 100
Separates “we found no problems” from “the submitted material did not test the problems.”
Evidence Density0-100
evidence_density = criteria with source evidence or verified absence / 50 * 100
Measures whether the source actually supports the assessment, not merely whether the pipeline returned data.
Domain BreakdownA-E scores
domain_score[batch] = sum of 5 maturity scores in that domain (0-15)
Shows whether readiness is concentrated in operating model, platform, governance/value, data, or service architecture.

Phase 3 Output Contract

Phase 3 produces structured JSON consumed by the dashboard. It cannot fabricate optimism because the visual scorecard, classification, evidence density, domain scores, anti-pattern findings, and source gaps are locked by deterministic validation.

{
  "phase_3_strategy": {
    "executive_summaries": {
      "transformation_lead": "operating model, decision rhythm, learning, portfolio, and adoption lens",
      "technology_lead": "architecture, platform, lifecycle, observability, safety, and data lens",
      "business_service_owner": "value stream, service ownership, customer impact, and capability lens"
    },
    "evidence_summary": "supported findings only, with adjacent context clearly labeled",
    "domain_diagnosis": {
      "A": "Adaptive Operating Model",
      "B": "Enterprise AI Architecture & Platform Readiness",
      "C": "AI Strategy, Governance & Value Realization",
      "D": "Data Foundations, Ownership & Accessibility",
      "E": "Business Capability & Service Architecture"
    },
    "planning_decision": "GO | WARN | BLOCK | NO_GO",
    "findings_mode": "used when confidence is LOW",
    "remediation_roadmap": [
      "Kickstart: diagnose, select domain, run vertical slice, validate value and safety",
      "Building the System: integrate architecture, platform, data, teams, rhythm, governance, and learning"
    ]
  }
}

Kickstart

Early roadmap phases diagnose readiness, select a service-area domain, validate a bounded vertical slice, test value and safety, and convert pilot learning into a playbook. This is where Impact Statements, Launch-and-Learn, Secure-by-Design, and evidence-backed business cases become practical.

Building the System

Later roadmap phases integrate architecture, platform, data products, service teams, operating rhythm, embedded governance, AI lifecycle management, Sense & Respond loops, and scaled learning across value streams.

The Report Is Checked After It Is Written

The engine treats generated strategy as another artifact to verify. The first complete report draft can still be rejected, rewritten, or blocked if it exceeds the evidence, leaks internal reference material, misuses a tactic, or turns uncertainty into advice.

Fact-check splitsummary and roadmap
Narrative claims and roadmap actions are checked separately so a good summary cannot hide an unsupported plan.
Regeneration loopbounded retry
Unsupported claims can trigger a narrower regeneration pass. Repeated weak grounding escalates to sanitation or quality-gate blocking.
Tactic validationvalid IDs only
Roadmap actions cite tactic IDs only when the validated finding matches an approved pattern. Generic validation work may intentionally have no tactic ID.
Strategy sanitationclaim hygiene
Unsupported fragments are rewritten, removed, or quarantined. The hygiene appendix keeps the decision visible without promoting unsafe advice.
Privacy readabilityfluent anonymization
Customer-facing prose is rewritten into neutral language such as “the assessed organization” and “the service area” instead of leaking names or filling the report with placeholders.
Quality gateGO / WARN / BLOCK
The final gate considers evidence density, delivery integrity, confidence bracket, fact-check support, sanitation state, and source coverage gaps.

Real Model Roles and Fallback Chains

Model routing is centralized in src/models.ts. The production configuration uses stronger reasoning models for forensic audit, targeted rescan, synthesis escalation, roadmap creation, fact-checking, and quality-gate explanation. Fallbacks are ordered by task shape: fast safety checks can start with Gemini Flash, independent verification uses Gemini Pro-class models, and high-stakes synthesis routes through Claude Sonnet/Opus and GPT-5.5 profiles.

Step / TaskPrimary ModelFallback Chain
Security pre-flight / DLP scangemini-3.5-flashgemini-2.5-progpt-5.5
Phase 1 forensic audit, batches A-Eclaude-sonnet-4-6gpt-5.5 audit profile → gemini-2.5-pro
Targeted rescan for weak evidenceclaude-opus-4-7gpt-5.5 roadmap profile → claude-sonnet-4-6gemini-2.5-pro
Independent evidence checkgemini-3.1-pro-previewclaude-sonnet-4-6gpt-5.5 evidence-check profile
Anti-pattern adjudicationgpt-5.5 fact-check profileclaude-opus-4-7
Evidence summary / diagnosisclaude-sonnet-4-6gpt-5.5 synthesis profile → gemini-2.5-pro
Phase 3 escalation modeclaude-opus-4-7gpt-5.5 escalation profile → gemini-3.1-pro-preview
Roadmap synthesisclaude-opus-4-7gpt-5.5 roadmap profile → gemini-3.1-pro-preview
Fact-checkgpt-5.5claude-sonnet-4-6gemini-2.5-pro
High fact-check retrygpt-5.5 high reasoningclaude-sonnet-4-6
Quality Gate explanationgpt-5.5 quality-gate profileclaude-sonnet-4-6

The Difference Between Tool Adoption and Value Creation

A generic AI-readiness survey asks whether the organization has a strategy, tools, pilots, data, and governance. The AI Transformation Engine asks whether the documents prove that those things operate together as a value-creation system.

Tool Adoption Without Readinessfalse positive
A chatbot rollout, model gateway, or pilot portfolio can look advanced while decision rights, data context, release control, service ownership, and value measurement are still missing.
Governance Without Enablementcontrol theater
Policies and approvals do not prove safe value creation unless controls are embedded into workflows, architecture, data access, release lifecycle, and escalation paths.
Pilots Without Absorptionpilot purgatory
A pilot is useful only when it tests value, safety, integration, adoption, learning, and the transition from Kickstart to Building the System.
Data Without Meaningcontext failure
Data becomes fuel only when it carries semantic and service context. Raw availability is not the same as usable AI grounding.
Architecture Without Ownershipscaling friction
AI cannot scale safely when service boundaries, ownership, platform capabilities, and human accountability are unclear.
The differentiator is not “AI readiness.” The differentiator is evidence-gated proof that the organization has the operating, architectural, data, governance, and service structures required for AI to create measurable value safely.