AI Transformation Assessment Engine — Evidence-Gated Scanner Architecture
This engine is not an AI tool checklist. It is a structured analytical pipeline for testing whether an organization can sense, decide, design, deliver, govern, learn, and scale AI-enabled value creation safely. AI readiness is organizational and architectural readiness, not tool readiness.
The engine evaluates actual operating evidence: decision rights, service-area ownership, brownfield integration, AI lifecycle control, semantic data context, value framing, embedded governance, safety controls, learning loops, and business architecture. Plans and slogans score low. Tool adoption alone does not prove readiness. Pilots become meaningful only when they are tied to value, safety, learning, and repeatable scaling.
The architecture is intentionally forensic. It uses local parsing, parse-quality checks, DLP scanning, five parallel audit batches, independent evidence verification, targeted rescans, deterministic metrics, confidence-gated synthesis, fact-checking, privacy-aware sanitation, and a GO/WARN/BLOCK quality gate.
The Forensic Analysis Pipeline
The pipeline separates extraction, verification, scoring, synthesis, and validation. The scanner extracts candidate evidence. The verifier checks whether the source material really supports it. The mathematics layer scores only verified data. The synthesizer receives locked metrics as boundaries. The fact-checker and sanitation layer then remove unsupported or privacy-unsafe claims before the report is shown.
Why This Order Matters
Extraction is separated from recommendation
The forensic auditor cannot inflate scores because a future roadmap would sound attractive.
Verification happens before scoring
Unsupported claims are downgraded or rescanned before deterministic metrics can see them.
Mathematics caps optimism
Readiness is bounded by evidence density, anti-pattern coverage, and delivery integrity.
Reference KB is not customer proof
The Knowledge Base can shape rubric interpretation and tactics, but it cannot prove current state.
The synthesizer cannot rewrite diagnosis
Phase 3 receives locked scores, gaps, and confidence bracket as hard boundaries.
The engine can refuse to plan
LOW-confidence runs are routed to findings mode rather than polished but unsafe advice.
Agentic Verification: The Engine Decides What the Evidence Allows
The engine is agentic within guardrails. It does not simply pass text from one model call to the next. It evaluates intermediate results, compares them against the source, and makes bounded process decisions about what should happen next. Those decisions are not open-ended autonomy; they are evidence-gated controls inside the assessment workflow.
Three Roles. Three Constraints.
The engine behaves like a small assessment team. Each role has a bounded mandate, a different stage responsibility, and a different failure mode. The roles are not decorative prompt personas; they are control surfaces that keep parsing, scoring, synthesis, and validation separate.
AI Transformation Signal Scan
[1] Core AI Transformation Vocabulary
AI operating model, AI governance, AI lifecycle,
agent, agentic workflow, autonomous decision,
human-in-the-loop, prompt, model gateway,
evaluation, groundedness, hallucination, RAG,
retrieval quality, red teaming, guardrail,
AI value hypothesis, Impact Statement,
Launch-and-Learn, service area ownership
[2] Architecture, Data, and Safety Signals
API boundary, brownfield integration, service
boundary, model registry, prompt registry,
rollback, observability, drift, lineage,
semantic context, data product, data access,
privacy, masking, purpose limitation, Agent
Behavioral Contract, Secure-by-Design
[3] Value-Creation and Operating Signals
service blueprint, value stream, capability map,
service catalog, decision rights, portfolio
learning, kill / continue / pivot, business case,
value realization, adoption, learning loop,
Sense & Respond, platform-as-product,
Kickstart, Building the System
Keyword signal estimates likely source richness; it does not score readiness. Generic operating-model documents can still be valuable when they prove decision rights, service ownership, value streams, or governance mechanisms that directly connect to AI Transformation criteria.
Behavioural Rules
- You are an AI Transformation Forensic Auditor. Your job is to extract explicit proof of AI-enabled value-creation readiness from uploaded source material.
- Do not give the benefit of the doubt. Plans, slogans, aspirations, and future-tense claims score 1 at most.
- Tool presence is not readiness. A chatbot, model gateway, platform, RAG index, or pilot must show operating proof before it can support higher scores.
- General IT/process maturity is adjacent context unless it directly proves AI operating rhythm, AI architecture, data readiness, governance, safety, value realization, or service-area AI ownership.
- Silence is data. If the source is silent on a criterion, score it 0 and record the missing evidence instead of balancing the audit.
- Every captured evidence quote must be tagged as Policy, Process, Operational, Automation, Accountability, Architecture, Data, Governance, Learning, or Value.
- Anti-pattern absence is not automatically positive. It must be classified as confirmed present, partial finding, tested absent, or not assessed.
- Scanner output is provisional. Evidence-check can downgrade scores and request targeted rescans when raw material does not support the forwarded finding.
Behavioural Rules
- Source of truth: every diagnostic claim must trace to Phase 2 metrics, evidence-checked findings, or uploaded source material.
- Do not invent owners, scores, numbers, readiness level, business value, adoption, controls, or maturity.
- LOW confidence means findings mode: no directive roadmap, no tactic prescriptions, planning decision NO_GO, and explicit missing-evidence guidance.
- Use the Knowledge Base only as rubric, false-positive, risk-framing, and tactic context. It must never become current-state proof, evidence quote, or visible source reference.
- Write in fluent anonymized prose from the start. Do not mention company names, organization names, Knowledge Base document names, source labels, or placeholder redaction tokens.
- Roadmap actions must follow Kickstart then Building the System: diagnose, select domain, run vertical slice, validate value/safety, then integrate architecture, platform, data, teams, rhythm, governance, and learning.
- Every prescribed tactic action cites a valid tactic ID only when the locked finding matches the tactic pattern. Generic evidence-gathering actions may omit tactic IDs.
- Persona summaries must agree on facts; they differ only in lens, vocabulary, and emphasis.
Two Streams. One Diagnosis.
Every domain has two opposing streams. Stream A measures maturity: evidence of capability, ownership, operating proof, architecture, governance, data quality, learning, and value realization. Stream B measures anti-patterns: evidence of friction, fragility, decision fog, brittle integration, context-poor data, unsafe autonomy, or tool adoption without value proof.
High Maturity + Low Anti-Patterns
Genuinely adaptive only when anti-pattern coverage is high and absence has been meaningfully tested.
High Maturity + High Anti-Patterns
Structured but inconsistent. The organization has mechanisms, but friction blocks scale or trust.
Low Maturity + High Anti-Patterns
Transformation work is needed before AI-enabled value creation can scale safely.
Low Maturity + Low Anti-Patterns
Usually insufficient evidence unless the absence of anti-patterns was directly tested by source material.
Anti-pattern absence is positive only when verified. A zero anti-pattern score can mean tested absent, which is useful evidence of clearance, or not assessed, where the source simply did not cover the issue. Low burden with low coverage means unknown, not excellent.
The 50-Criteria AI Transformation Knowledge Base
Five domains. Each domain holds 5 maturity indicators paired against 5 anti-patterns. Each criterion carries 3 evidence questions. The result is a 50-criterion, 150-question scanner for whether AI can become safe, measurable, service-owned value creation.
Adaptive Operating Model
Maturity Indicators
- Is there a documented AI operating rhythm?
- Are decision rights assigned across business, IT, data, architecture, security, legal, and service areas?
- Are decisions traceable through cadence, approval records, backlog movement, or escalation logs?
- Are demands classified as predictable, standardized, iterative, exploratory, or transformative?
- Are different delivery paths used for different demand types?
- Is routing connected to capacity, competence, service ownership, value, cost profile, and risk level?
- Does transformation start from a safe experiment or vertical slice?
- Are pilot learnings used to refine operating model, architecture, data, and governance?
- Is there a transition from Kickstart validation to Building-the-System scaling?
- Are competence structures used to spread AI learning?
- Are pilots, failures, red-team results, and incidents converted into practices?
- Do business, AI, data, platform, and service teams learn together?
- Are initiatives designed to reallocate human capacity toward higher-value work?
- Are roles, handoffs, exception paths, and review points redesigned?
- Are employee experience, cognitive load, hidden review or rework cost, quality, and customer impact measured?
Anti-Patterns
- Are AI initiatives delayed because nobody clearly owns decisions?
- Do separate forums create repeated approvals instead of flow?
- Are ordinary design decisions escalated because boundaries are unclear?
- Are all AI initiatives forced through one delivery model?
- Are exploratory cases managed with rigid plans that assume known requirements?
- Are predictable automation and uncertain agentic work mixed without value, risk, or cost-profile routing logic?
- Do pilots remain disconnected from production, ownership, or outcomes?
- Are pilots called successful without integration, safety, adoption, or value proof?
- Does each use case start from scratch instead of reusing patterns?
- Is knowledge concentrated in a few experts or isolated teams?
- Are mistakes repeated because learning is not institutionalized?
- Do teams depend on individual heroics instead of shared practices?
- Is AI used mainly to speed fragmented tasks?
- Does AI create extra checking, rework, hidden review cost, hidden coordination, or workslop?
- Are people measured on activity rather than impact, learning, quality, and customer value?
Enterprise AI Architecture & Platform Readiness
Maturity Indicators
- Are integrations built through APIs, service boundaries, events, or governed connectors?
- Are legacy systems wrapped before agents depend on them?
- Are ownership, versioning, dependencies, and failure modes documented?
- Are models, prompts, agents, tools, datasets, configurations, and evaluations versioned?
- Is there a build-test-deploy-monitor-rollback lifecycle?
- Can production behavior be traced to model, prompt, data, tool, and release version?
- Are quality, safety, token/model spend, latency, routing performance, retrieval cost, cost-per-output, groundedness, drift, and business impact monitored?
- Are offline and online evaluation methods defined?
- Do alerts trigger when performance, data freshness, value, cost, or safety thresholds degrade?
- Are prompt injection, leakage, toxic content, unsafe tool use, and adversarial scenarios tested?
- Are guardrails, filters, access controls, and escalation paths embedded?
- Is there an Agent Behavioral Contract defining forbidden behaviors?
- Are shared AI services, model gateways, routing, caching, quota controls, budget alerts, data patterns, and templates offered as platform products?
- Are platform teams accountable to service-area teams as internal customers?
- Are adoption, flow, reliability, cost-aware developer experience, and value enablement measured?
Anti-Patterns
- Do use cases depend on manual exports, screen scraping, or fragile connectors?
- Does changing one system break another because boundaries are unclear?
- Are integrations created project-by-project with no reusable architecture?
- Are models, prompts, or agents moved into production without controls?
- Are prompt changes made manually without review or rollback?
- Is it impossible to reproduce why an AI system behaved a certain way?
- Are failures discovered mainly through complaints or business damage?
- Are there no baseline metrics for accuracy, token/model spend, cost-per-output, value, safety, or reliability?
- Are drift, retrieval degradation, hallucinations, cost anomalies, or unsafe outputs not systematically detected?
- Are guardrails trusted without test evidence?
- Can agents access tools, systems, or data without clear boundaries?
- Are safety reviews late gates instead of built into service design?
- Does every project build its own model, embedding, vector-store, retrieval, and monitoring stack?
- Are teams buying disconnected tools without standards, cost controls, budget alerts, or observability?
- Does fragmentation create duplicate AI spend, maintenance debt, and reduced innovation capacity?
AI Strategy, Governance & Value Realization
Maturity Indicators
- Does the organization define why AI matters to strategy, customers, values, and business model?
- Are ambition areas connected to service areas, value streams, or strategic domains?
- Are where-not-to-use boundaries explicit?
- Are initiatives started from explicit Impact Statements rather than tool ideas?
- Are outcomes defined through customer value, profitability, quality, flow, experience, or cost per workflow, proposal, ticket, conversation, document, assessment, or decision?
- Are value hypotheses tested and updated after pilots or releases?
- Are governance responsibilities embedded into service areas, platform, security, data, and business ownership?
- Does governance enable safe value creation rather than only blocking risk?
- Are exceptions, approvals, risks, and controls auditable?
- Are use cases classified by risk, autonomy, user impact, data sensitivity, and regulation?
- Are oversight, escalation, disclosure, logging, and accountability tied to risk?
- Are high-risk or autonomous scenarios blocked unless minimum controls are proven?
- Are investments managed through impact, risk, readiness, budgeting, forecasting, spend guardrails, cost, and learning logic?
- Are kill, continue, scale, and pivot decisions based on value-vs-cost evidence?
- Does pilot learning reprioritize the portfolio?
Anti-Patterns
- Is AI described generically without strategic choices?
- Is everything labeled an AI priority without prioritization logic?
- Are tools adopted before value is defined?
- Are initiatives selected because they are visible or technically interesting?
- Are savings, ROI, TCO, token/model spend, or efficiency claims made without baseline and post-measurement?
- Are spend, usage, or successful pilots celebrated without enterprise-level impact?
- Does governance slow flow without improving safety?
- Are teams unclear how to proceed safely, causing bypass or shadow AI?
- Are governance decisions disconnected from impact, architecture, and operations?
- Are use cases launched without risk classification?
- Is accountability unclear when AI-assisted decisions cause harm or failure?
- Are override, appeal, fallback, or escalation mechanisms missing?
- Does the backlog, platform usage, or AI spend grow without capacity, readiness, or value proof?
- Are weak initiatives kept alive by sunk cost, politics, or enthusiasm?
- Are investments focused on platform build-out, premium model usage, or tool expansion without near-term capability validation?
Data Foundations, Ownership & Accessibility
Maturity Indicators
- Are critical AI datasets owned by the service area or domain that understands context?
- Are owners accountable for quality, meaning, lifecycle, and usability?
- Are data products published with context, definitions, access rules, and usage expectations?
- Does data include semantic context, source meaning, process linkage, and operating conditions?
- Can AI systems understand what values mean in service context?
- Are critical data elements mapped to decisions, triggers, workflows, or outcomes?
- Are critical datasets cataloged, versioned, quality-checked, and traceable?
- Are freshness, completeness, correctness, and drift monitored?
- Can training, retrieval, inference, and decision data be traced after the fact?
- Are classification, purpose limitation, access rights, and privacy controls documented?
- Is access role-based, auditable, and aligned with ownership?
- Are masking, anonymization, minimization, retention, and boundary controls used where required?
- Are structured, unstructured, real-time, batch, feature, and retrieval patterns defined?
- Are RAG, embedding, search, feature-store, vector-store, and event patterns governed and reusable?
- Is retrieval quality, coverage, freshness, grounding, context growth, duplication, and retrieval cost measured?
Anti-Patterns
- Is data ownership centralized in IT or undefined, causing context loss?
- Are AI teams forced to interpret datasets without domain meaning?
- Do data issues stall AI because nobody owns correction or interpretation?
- Is data collected because available rather than because it drives decisions?
- Do models receive raw signals without business, process, or operational context?
- Are outputs unreliable because the model lacks meaning and decision context?
- Are outputs based on data whose origin, version, or quality cannot be proven?
- Are stale copies, manual files, or undocumented transformations used?
- Are quality failures found only after model failure or escalation?
- Can tools access sensitive data without purpose, logging, or approval?
- Are copies moved into experimentation without lifecycle control?
- Are privacy and security controls added late?
- Does each team create separate embeddings, vector stores, or knowledge bases without ownership or cost visibility?
- Do RAG systems use stale, incomplete, or unverified sources?
- Are retrieval failures, unbounded context growth, and duplicated retrieval costs treated only as model failures?
Business Capability & Service Architecture
Maturity Indicators
- Is there a current service catalog, capability map, or equivalent model?
- Are opportunities linked to services, capabilities, customer paths, or value streams?
- Are dependencies between capabilities, platforms, data, and outcomes visible?
- Are target services mapped across customer journey, frontstage, backstage, systems, data, and handoffs?
- Are bottlenecks, decisions, failure modes, and automation candidates visible?
- Are interventions placed where they improve the whole value stream?
- Are features decomposed into need, capability, component, data, platform, and work packages?
- Can each use case be traced to customer need, outcome, cost-to-serve where evidence exists, and service architecture?
- Are requirements, acceptance criteria, risks, and dependencies explicit before build?
- Are integrated teams accountable for business, application, data, and AI outcomes?
- Are platform teams positioned as enablers for service-area teams?
- Are AI solutions owned as living products rather than temporary projects?
- Is scaling sequenced by readiness, architecture, data, risk, and cost-to-serve readiness where evidence exists?
- Are pilot patterns converted into playbooks, templates, platform capabilities, and guardrails?
- Is continuous improvement built into rollout through feedback and reassessment?
Anti-Patterns
- Are use cases listed without connection to service catalog or capability map?
- Is it unclear which service area owns the change?
- Are dependencies discovered only during implementation?
- Is AI applied to isolated tasks without surrounding flow?
- Does automation speed one step while increasing downstream work?
- Are customer experience and end-to-end flow ignored?
- Are teams building AI components without linkage to capability or need?
- Are requirements vague, unstable, or tool-led?
- Is traceability between feature, data source, platform dependency, value outcome, and cost-to-serve missing?
- Are AI teams separate from the service areas that need to own outcomes?
- Are business, data science, application, and platform work sequential handoffs?
- Does ownership disappear after pilot completion or vendor delivery?
- Is rollout scaled before architecture, data ownership, cost-to-serve readiness, and operating model are ready?
- Are pilot learnings not converted into reusable patterns?
- Does every service area repeat mistakes independently?
0-3 Evidence Scale and AI Maturity Ladder
The scanner asks what the submitted material proves. A statement that “AI is strategic” is not enough. A pilot is not enough. A platform purchase is not enough. Higher scores require operating proof, measured feedback loops, tested controls, ownership records, or traceable production behavior.
No evidence found, or the source is not assessable for that criterion.
Intent, plans, pilots, future-tense claims, or isolated examples without operating proof.
Functioning mechanism, ownership, process, control, or delivery practice is described.
Mechanisms, measurement, enforcement, learning loops, and repeatable production behavior are visible.
Evidence Check Statuses
Readiness Classification
Evidence Categories
The score asks how strong the evidence is. The category tag asks what kind of evidence it is. Every evidence quote is tagged with exactly one category: Policy, Process, Operational, Automation, Accountability, Architecture, Data, Governance, Learning, or Value.
Source Rules and Knowledge Base Guardrails
The engine treats the uploaded source pack as untrusted customer evidence and the remote Knowledge Base as internal methodology context. Those two sources are never allowed to collapse into one another.
{
"source_evidence": "Only uploaded assessment material can prove current state.",
"reference_kb": "Rubric context, validation questions, false-positive checks, risk framing, and tactics only.",
"forbidden_kb_uses": [
"customer_current_state_claim",
"audited_finding",
"evidence_quote",
"score_justification_without_uploaded_source",
"visible_document_name",
"visible_company_or_organization_name"
],
"general_process_boundary": "Generic IT/process maturity is adjacent context unless it directly proves an AI Transformation criterion.",
"privacy_output_rule": "Write fluent anonymized prose from the start. Do not emit company names or placeholder tokens."
}
What to Submit
The assessment accepts any combination of documents. More document types provide richer signal across all five domains. Absence of entire document categories is itself a maturity signal. Do not manufacture missing documents for the assessment; absence is part of the diagnosis.
The Firewall Between AI and Roadmap
Phase 2 contains no generative AI. It applies deterministic arithmetic to evidence-checked audit logs. Unsupported scanner output has already been rescanned or downgraded before this layer runs.
Phase 3 Output Contract
Phase 3 produces structured JSON consumed by the dashboard. It cannot fabricate optimism because the visual scorecard, classification, evidence density, domain scores, anti-pattern findings, and source gaps are locked by deterministic validation.
{
"phase_3_strategy": {
"executive_summaries": {
"transformation_lead": "operating model, decision rhythm, learning, portfolio, and adoption lens",
"technology_lead": "architecture, platform, lifecycle, observability, safety, and data lens",
"business_service_owner": "value stream, service ownership, customer impact, and capability lens"
},
"evidence_summary": "supported findings only, with adjacent context clearly labeled",
"domain_diagnosis": {
"A": "Adaptive Operating Model",
"B": "Enterprise AI Architecture & Platform Readiness",
"C": "AI Strategy, Governance & Value Realization",
"D": "Data Foundations, Ownership & Accessibility",
"E": "Business Capability & Service Architecture"
},
"planning_decision": "GO | WARN | BLOCK | NO_GO",
"findings_mode": "used when confidence is LOW",
"remediation_roadmap": [
"Kickstart: diagnose, select domain, run vertical slice, validate value and safety",
"Building the System: integrate architecture, platform, data, teams, rhythm, governance, and learning"
]
}
}
Kickstart
Early roadmap phases diagnose readiness, select a service-area domain, validate a bounded vertical slice, test value and safety, and convert pilot learning into a playbook. This is where Impact Statements, Launch-and-Learn, Secure-by-Design, and evidence-backed business cases become practical.
Building the System
Later roadmap phases integrate architecture, platform, data products, service teams, operating rhythm, embedded governance, AI lifecycle management, Sense & Respond loops, and scaled learning across value streams.
The Report Is Checked After It Is Written
The engine treats generated strategy as another artifact to verify. The first complete report draft can still be rejected, rewritten, or blocked if it exceeds the evidence, leaks internal reference material, misuses a tactic, or turns uncertainty into advice.
Real Model Roles and Fallback Chains
Model routing is centralized in src/models.ts. The production configuration uses stronger reasoning models for forensic audit, targeted rescan, synthesis escalation, roadmap creation, fact-checking, and quality-gate explanation. Fallbacks are ordered by task shape: fast safety checks can start with Gemini Flash, independent verification uses Gemini Pro-class models, and high-stakes synthesis routes through Claude Sonnet/Opus and GPT-5.5 profiles.
| Step / Task | Primary Model | Fallback Chain |
|---|---|---|
| Security pre-flight / DLP scan | gemini-3.5-flash | gemini-2.5-pro → gpt-5.5 |
| Phase 1 forensic audit, batches A-E | claude-sonnet-4-6 | gpt-5.5 audit profile → gemini-2.5-pro |
| Targeted rescan for weak evidence | claude-opus-4-7 | gpt-5.5 roadmap profile → claude-sonnet-4-6 → gemini-2.5-pro |
| Independent evidence check | gemini-3.1-pro-preview | claude-sonnet-4-6 → gpt-5.5 evidence-check profile |
| Anti-pattern adjudication | gpt-5.5 fact-check profile | claude-opus-4-7 |
| Evidence summary / diagnosis | claude-sonnet-4-6 | gpt-5.5 synthesis profile → gemini-2.5-pro |
| Phase 3 escalation mode | claude-opus-4-7 | gpt-5.5 escalation profile → gemini-3.1-pro-preview |
| Roadmap synthesis | claude-opus-4-7 | gpt-5.5 roadmap profile → gemini-3.1-pro-preview |
| Fact-check | gpt-5.5 | claude-sonnet-4-6 → gemini-2.5-pro |
| High fact-check retry | gpt-5.5 high reasoning | claude-sonnet-4-6 |
| Quality Gate explanation | gpt-5.5 quality-gate profile | claude-sonnet-4-6 |
The Difference Between Tool Adoption and Value Creation
A generic AI-readiness survey asks whether the organization has a strategy, tools, pilots, data, and governance. The AI Transformation Engine asks whether the documents prove that those things operate together as a value-creation system.