Projects
Forward Deployed Engineering Reference architecture and operating-model analysis
/

Forward Deployed Engineering

A reference architecture and operating model for taking complex enterprise and AI systems from prototype to governed production — and an account of where the model breaks when nobody is watching for it.

Enterprise Architecture AI Engineering Platform Engineering Zero Trust DevSecOps AI Governance
Type
Architecture research & reference design
Domain
Enterprise AI and complex software delivery
Focus
Architecture · Security · Delivery · Governance
Status
Published reference architecture
Verified
9 August 2026

What this is. I wrote this reference architecture to work out how forward deployed engineering actually moves complex AI and enterprise systems from prototype into governed production, and where the operating model and the architecture supporting it tend to fail. It is authored research. It does not describe a client engagement, and no deployment, customer, saving or audit result is claimed anywhere on this page. The four scenarios in Reference scenarios are composites assembled from common industry constraints and are labelled as such.

The source material I started from contained several claims that did not survive verification, including a specific financial saving, a regulatory audit outcome and the assertion that a knowledge graph produces zero hallucinations. I have published those corrections rather than quietly dropping them. See Corrections.

01 The deployment gap

A vendor sells a capability. A customer needs an outcome. Between those two things sit five boundaries, and most implementations fail at one of them rather than at the technology itself.

Nothing about that is new. What has changed is the volume and the stakes: enterprises are now buying systems whose behaviour is probabilistic, whose integration surface reaches into their most sensitive data, and whose value depends entirely on people changing how they work. AWS announced a billion-dollar commitment to a dedicated forward deployed engineering organization in June 2026 on exactly this reasoning.[1] Accenture announced a Microsoft-focused forward deployed engineering practice three months earlier, on its own account.[5] OpenAI and Anthropic each spun delivery into a separate services venture rather than scaling it inside the lab.[6][7]

Select a boundary to see what actually goes wrong there.

Product
capability
Integration
Customer
systems
Real
workflows
Adoption
Production
outcomes

Each of these is solvable. None of them is solvable by documentation, and only the first is reliably solvable by a pre-sales architect who leaves after the design review. That is the argument for forward deployment, and it is worth stating in its weakest defensible form rather than its strongest: some problems sit in the seams between systems and teams, and the seams are where nobody has authority to change code.

02 What makes engineering “forward deployed”

Forward deployed engineering is an operating model, not an architecture. The architecture on this page is a synthesized reference design for supporting that model. There is no canonical forward deployed architecture, and I have found no standards body that defines one.

Five properties separate it from adjacent ways of working. An engagement missing several of them is doing something else, whatever it is called, and the label then obscures rather than clarifies what is being bought. Marty Cagan’s argument for the model is that it accelerates product discovery, and that it “applies much more broadly than just in the most challenging product situations” — worth weighing as advocacy from a product-management perspective rather than as evidence.[11]

01 Customer embedding

The engineer works close enough to users and systems to observe operational reality rather than requirements. Embedding is operational, not geographic. Palantir’s own posting for the role it originated caps travel at 25%,[2] and PostHog runs its forward deployed engineer fully remote and asynchronous.[12]

02 Engineering authority

The role can architect and write production software, not only recommend that someone else does. OpenAI describes its forward deployed engineers as owning “discovery, technical scoping, system design, build, and production rollout”.[3] Without the authority to build, the role becomes advisory work under a more energetic title.

03 End-to-end ownership

Discovery, scoping, architecture, implementation, evaluation, rollout and stabilization stay connected to the same people. Every handoff between those stages loses the context that made the previous stage correct.

04 Field-to-product feedback

Deployment problems that recur become inputs to platform and product work. This is the property most often claimed and least often instrumented, and its absence is what turns the model into custom consulting with a software company’s cost base.

05 Outcome orientation

Success is measured by production adoption and workflow effect, not by configuration completed or milestones closed. This is also where the model is most easily faked, because adoption metrics are easy to define generously.

What it is not

It is not a security discipline, a compliance mechanism or an architecture style. It is a way of organizing delivery. Everything security-relevant on this page comes from controls that would be required regardless of who was writing the code.

The label is contested, and that is worth knowing

The term expanded very quickly. Job postings for forward deployed roles were reported by the Financial Times to have risen more than 800% between January and September 2025, a figure repeated widely although the underlying dataset is not identified in any source I could find.[19] That growth has pulled in work that does not match the definition. Constellation Research’s Ray Wang put it bluntly: “Some of these FDEs don’t necessarily fit the definition and are glorified sales and customer success people.”[16] Gergely Orosz, reading a major cloud vendor’s posting for the role, translated it as “you are a contractor who codes at a customer’s office”.[13] Sierra’s own account concedes that the term “can mean so many things” that it risks becoming meaningless.[20] Even the encyclopaedia entry for the role is thin, resting largely on vendor self-description.[18]

The vendors are not consistent either. AWS positions its practice explicitly against traditional consulting.[1] Databricks files its AI forward deployed engineering team under Professional Services Operations and describes the work as “professional services engagements”.[8] Vercel’s posting says “this is not a traditional consulting role” while reporting the role to a Director of Professional Services.[9] None of that makes the model illegitimate. It does mean that “we use forward deployed engineers” carries almost no information on its own, and that an enterprise buying the model should ask what specifically it is buying.

03 FDE and adjacent roles

The boundaries between these roles vary by company. What follows is a comparison of common operating patterns, not a set of formal industry definitions, and any given organization will place at least one of them differently.

Show all seven roles side by side

Scrollable on narrow screens. The single-role view above carries the same content in a form that reads on a phone.

Comparison of forward deployed engineering with six adjacent roles across nine dimensions
DimensionFDESolutions ArchitectSolutions EngineerProfessional ServicesCustomer EngineerProduct EngineerTAM
Primary objectiveProduction outcome in one environmentA defensible designTechnical validation of a dealBounded delivery of a known patternTechnical adoption and enablementCapability for many customersSustained account health
Lifecycle stagePost-sale through operationsPre-sale and early designPre-salePost-sale, fixed scopePre- and post-saleContinuous roadmapOngoing
Customer proximityEmbedded in the workflowPeriodic, workshop-basedMeeting-basedProject-basedRegular, advisoryIndirectRelationship-based
Coding depthProduction code, substantialReference material, prototypesDemos and proofs of conceptConfiguration and integration codeSamples and acceleratorsProduct codeLittle to none
Production accountabilityYes, through stabilizationNoNoTo acceptance criteriaShared, informalFor the platform, not the deploymentEscalation ownership
CustomizationDeep, environment-specificDesign-level onlyIllustrativeWithin a defined catalogueLightNone — generalizes insteadNone
Product feedback dutyExplicit and instrumentedInformalCompetitive and feature gapsDelivery frictionAdoption blockersReceives itAccount-level themes
Operational ownershipUntil handoff is provenNoneNoneUntil acceptanceNonePlatform SLOsCoordinates, does not operate
Typical durationMonths, occasionally longerWeeksDays to weeksFixed-term projectOngoing, part-timeContinuousContract lifetime

On the titles themselves. “Forward Deployed Software Engineer” and “Forward Deployed Security Engineer” are both attested at named employers.[2][3] “Technical Deployment Lead” appears at OpenAI, always compounded with the forward deployed engineering function rather than standing alone.[3] I could not attest “Forward Deployed Platform Engineer” as a title in use at any employer, so I am not presenting it as one; the nearest real examples are Sierra’s Forward Deployed Infrastructure Engineer and Cognition’s Deployed Engineer, which is not the same phrase. “Customer Engineer” predates this wave as a pre-sales title at Google Cloud and is not a synonym.

04 Forward-deployed reference architecture

Six runtime layers with two concerns that cut across all of them. This is a reference design I assembled to reason about the model, not a description of any vendor’s product. Named technologies are examples of what can occupy a slot, never requirements.

Forward-deployed reference architecture Six stacked runtime layers on the left, from Experience and operational workflow at the top down through AI and agent runtime, domain and semantic layer, data and integration plane, application and platform runtime, and software delivery and supply chain at the bottom. Two vertical bands on the right span the full height: identity, security and governance; and observability and operations. RUNTIME STACK CROSS-CUTTING Identity, security & governance Observability & operations Experience & operational workflow Applications, dashboards, embedded workflow surfaces, approval interfaces, downstream APIs AI & agent runtime Model access, routing, context assembly, retrieval, tools, orchestration, guardrails, evaluation Domain / semantic layer Domain objects, relationships, business semantics, lineage, policy-aware metadata, actions Data & integration plane APIs, events, batch, change data capture, stores, transformation, legacy and OT adapters Application & platform runtime Compute, service identity, configuration, secrets, feature flags, service-to-service paths Software delivery & supply chain Source control, CI/CD, infrastructure and policy as code, signing, SBOM, promotion, rollback Every layer is subject to both cross-cutting bands. A control that exists in only one layer is not a control.
Figure 1 — The runtime stack and the two concerns that cross it. Security and observability are drawn as bands rather than layers because treating either as a layer is how they end up owned by nobody.

Select a layer for its purpose, the split of responsibility between vendor and customer, what tends to go wrong there, and example technologies.

What this diagram deliberately does not require. Kubernetes, a knowledge graph, an ontology, microservices, edge computing, serverless, GraphRAG, a data mesh, a particular cloud and a particular model vendor are all absent from the layer definitions. Each is an implementation choice that some deployments justify and most do not. Start from the capability and the constraint; introduce the technology when you can say what it is for.

05 One operating model, different deployment boundaries

The operating model barely changes across these five. What changes is where the trust boundary falls, and every consequential decision on the page follows from that: how software is updated, how the engineer reaches the system, whether telemetry ever arrives, and who holds the credentials.

Where the trust boundary falls in each deployment topology Five components — control plane, application runtime, data plane, identity, and telemetry and update path — positioned along a horizontal axis running from the customer estate on the left to the vendor estate on the right. A vertical dashed line marks the trust boundary and moves as the selected topology changes. The same information is written out as a component-placement list at the top of the panel below the diagram. CUSTOMER ESTATE VENDOR ESTATE TRUST BOUNDARY Control plane Control plane Application runtime Application runtime Data plane Data plane Identity Identity Telemetry & updates Telemetry & update path
Figure 2 — The same five components, repositioned. Components left of the dashed line are governed by the customer’s controls; components to its right by the vendor’s. Side is also stated in the panel below; position is authoritative.

No platform has to support all five. Supporting the air-gapped case imposes design constraints on everything else — offline licensing, signed offline bundles, deferred telemetry, local policy decision points, an update path that survives months of disconnection. Building those for a customer base that is entirely cloud-resident is a large, permanent tax on velocity. Deciding which topologies you will not support is an architecture decision and should be recorded as one.

06 Decisions that matter

These are the ten decisions I have found determine whether a forward deployed engagement stays maintainable. Each has a genuine case on both sides; the failure is not choosing wrong, it is choosing implicitly and discovering the choice two years later during an upgrade.

Centralized data versus federated query Data

Drivers. Data residency and sovereignty rules, cross-border transfer mechanisms, the sensitivity of the underlying records, query latency tolerance, and whether the analytical questions are known in advance.

Centralize when the analytical workload is exploratory, the data can lawfully move, and join performance across the whole corpus matters more than locality.

Federate when raw records cannot cross a jurisdictional or organizational boundary. Push the query to the data, apply local masking and aggregation, and return only results. The cost is real: federated queries are slower, harder to optimize, harder to debug, and each local node becomes a component you have to operate and patch.

Risk if implicit. Teams build a central lake, then discover a residency constraint, then bolt on regional exceptions until the model is neither central nor federated and nobody can describe where a given record lives.

Cloud inference versus local inference AI

Drivers. Data sensitivity, latency budget, connectivity, unit cost at expected volume, model capability required, and how often the model needs replacing.

Cloud when you need frontier capability, volumes are variable, and the data can leave the estate under an acceptable contractual and technical control set.

Local when connectivity is constrained or absent, latency is measured in tens of milliseconds, or the data cannot leave. Accept the consequences: a smaller model, a deployment pipeline for model artifacts, drift monitoring you have to build, and a hardware refresh problem.

The middle path that usually wins is decoupling training from inference — train centrally, ship a compact evaluated artifact to the edge, run drift detection locally, and queue retraining requests for whenever the link returns.

Product configuration versus custom extension Delivery

Drivers. How far the requirement is from the product’s intended use, how many other customers would want it, who will maintain it in three years, and whether the platform has an extension contract worth depending on.

Configure when the gap is a setting, a policy or a template. Configuration survives upgrades; code does not, unless someone maintains it.

Extend when the requirement is genuinely specific and the extension point is versioned and tested. Register the extension, name its owner, record which platform API version it depends on, and set a disposition date.

Risk if implicit. Custom code written under delivery pressure with no owner and no registry entry is, in my reading, the most consequential source of upgrade paralysis in this model.

Synchronous versus asynchronous integration Integration

Drivers. Whether the caller can wait, the availability of the downstream system, whether the operation is idempotent, and how failures should surface to a human.

Synchronous when the user is waiting for a decision and the downstream system has a credible availability target. Bound the timeout, define the fallback, and decide in advance what the interface says when the dependency is down.

Asynchronous when the downstream is a legacy system with batch semantics, the operation is expensive, or partial availability is normal. The cost is that you now own a queue, a retry policy, a dead-letter path, an idempotency key strategy and a reconciliation process.

Retrieval over documents versus retrieval over a graph AI

Drivers. Whether the questions are relationship-heavy, whether the entities are resolvable, whether provenance must be traceable to a specific record, and whether anyone will maintain the schema.

Document retrieval when the corpus is textual, the questions are local, and the answer lives in a passage. It is cheaper, simpler and has a shorter failure chain.

Graph retrieval when answering requires traversing relationships between resolved entities, when authorization needs to be expressed over those relationships, or when provenance must point at identified nodes and edges rather than at a chunk of text.

What neither buys you. Correctness. See GraphRAG and its limits for the measured numbers, which are worse than most architecture diagrams imply.

Direct model access versus an AI gateway AI

Drivers. Number of applications, need for central policy, cost attribution, the likelihood of changing model vendor, and whether anyone needs a single audit trail of model use.

Direct when there is one application, one model and one team. A gateway for a single caller is infrastructure you maintain for no benefit.

Gateway when you need central authentication, per-tenant quota and cost attribution, prompt and model version pinning, uniform logging, and the ability to substitute a model without touching applications. The gateway becomes a dependency on the critical path and needs its own availability target and failure mode.

Autonomous agent action versus human-approved action AI

Drivers. Reversibility of the action, blast radius, the cost of a false positive against the cost of delay, and how many approvals a reviewer will see per day.

Autonomous when the action is reversible, bounded and cheap to get wrong — drafting, classifying, enriching, proposing.

Human-approved when the action moves money, changes access, deletes data, contacts a customer or cannot be undone. Approval is a real control that degrades under volume, so it has to be rationed. Requiring approval for everything produces a queue that gets cleared, not read.[75] Automation-induced complacency is a long-established human-factors finding rather than a new observation about AI.[76]

The control that does not depend on attention is authorization enforced in the downstream system. OWASP’s own guidance is explicit: “implement authorization in downstream systems rather than relying on an LLM to decide if an action is allowed”.[43]

Shared vendor plane versus customer-isolated runtime Platform

Drivers. Regulatory posture, tolerance for shared-tenancy risk, the customer’s own security review process, and what the vendor can afford to operate.

Shared when the customer accepts multi-tenant isolation and the vendor can demonstrate it. Operating cost per customer stays low and upgrades are uniform.

Isolated when a regulator, a contract or a threat model requires it. Every isolated runtime is a separate environment to patch, monitor, back up and prove. Ten isolated runtimes multiply the operational surface tenfold; revenue rarely follows the same curve.

Standing remote support access versus customer-controlled just-in-time access Security

Drivers. Incident response time, the customer’s privileged access model, audit requirements, and whether the vendor is treated as a third party or a sub-processor.

Standing access shortens mean time to repair and is what most engagements drift into. It also means a vendor credential compromise is a customer production incident, and it is the finding an assessor will write up first.

Just-in-time access costs minutes at the start of every incident and removes an entire category of standing risk. Requested through the customer’s own workflow, scoped to a role, time-boxed, session-recorded, reviewed. Design for this from the first day, because retrofitting it after two years of standing access is an organizational fight, not a technical one.

Customer-specific code versus product primitive Product

Drivers. How many customers need it, whether it exposes a missing platform capability, who will own its lifecycle, and whether it can be tested and versioned independently.

Keep it specific when the requirement encodes one customer’s process. Promoting local process into a shared platform makes every future customer carry it.

Promote to a primitive when the same shape has appeared three times, the abstraction is stable, and moving it into the core reduces total delivery cost across the portfolio rather than just in the next engagement.

The intermediate step people skip is an extension contract — a supported, versioned point where the specific thing can live without becoming core. See Productization economics.

07 Engagement lifecycle

Eleven stages. The first one is the one most organizations skip, and skipping it is why the model gets a reputation for being expensive.

Deliberately without a calendar. The source material I worked from offered two adoption timelines for structurally similar programmes that differed by a factor of two to three, neither attributed. Published month ranges for work like this are estimates that a sponsor will hold you to. What is durable is the ordering and the exit criteria: qualification before discovery, evaluation before productionization, readiness proven before rollout, handoff proven before the engagement is called finished.

08 The feedback loop

This is the part of the model that determines whether it is a business or a service line with unusually expensive staff.

The field-to-platform feedback loop A cycle beginning at customer workflow, producing a field signal, which reaches the delivery or forward deployed engineering function. From there it branches to product engineering and to research or model teams. Both branches converge on a platform change, which feeds the next deployment, which returns to customer workflow. A dashed return path runs from delivery back to customer workflow, bypassing the platform change entirely; it is labelled as what happens when no platform change is made. Customer workflow Field signal Delivery / forward deployed Product engineering Research / model Platform change Next deployment Bypass: bespoke every time
Figure 3 — The loop, and the dashed path that replaces it when field learning never reaches the platform. The dashed route still delivers working systems; it just delivers them at constant marginal cost forever.

The model scales only when field learning improves the platform. If every engagement stays bespoke, forward deployment is custom consulting carrying a software company’s valuation. F-Prime Capital, writing from an investor’s perspective, proposes a working threshold: if more than 30–40% of deployments require significant forward deployed effort, the problem has stopped being go-to-market and become product design.[15] That figure is a practitioner heuristic rather than a measured finding, and I would treat it as a prompt to instrument rather than a number to manage to. The underlying claim is harder to argue with: “if every deal requires bespoke engineering, you don’t have a product; you have a consultancy with a logo”.[15]

There is a serious counter-position. Andreessen Horowitz argues that optimizing for gross margin percentage is the wrong objective early, and points at ServiceNow and Workday, both of which entered public markets in the fifties and low sixties and reached the mid-to-high seventies years later.[17] Both parties are investors and both have a position to talk. The reconciliation I find defensible is narrow: services-heavy delivery is a legitimate investment when it is buying a repeatable capability, and it is a permanent cost when it is buying a customer. The instrumentation that tells you which one is happening is in Measuring outcomes, and most organizations running this model do not have it.

09 Productization economics

Every forward deployed engagement produces code that should not stay where it was written. The discipline is knowing which code that is, and having somewhere for it to go that is not the core product.

The productization progression A seven-stage progression from customer request, to one-off solution, to repeated field pattern, to reusable accelerator, to platform abstraction, to product capability, to self-service. A note marks that most organizations stop at stage two and repeat it. Stage two is outlined in red, stage three in amber and the last two in green, matching the progression from one-off work to reusable capability. Customerrequest One-offsolution Repeatedfield pattern Reusableaccelerator Platformabstraction Productcapability Self-service where most engagements loop reuse begins here
Figure 4 — The progression, and the loop that replaces it. Moving from stage two to stage three requires only that someone notices the repeat; moving from three to four requires that someone is funded to.

The promotion questions

Before customer-specific work is promoted toward the core, these nine questions decide whether it should be. A “no” to the last one is disqualifying regardless of the others.

  • How many customers actually need this, as opposed to how many might?
  • Is the requirement domain-specific, or does it encode one organization’s internal process?
  • Does it expose a platform primitive that is missing, or does it work around one that exists?
  • What permanent complexity does supporting it add, and to whom?
  • Who owns its lifecycle after the engagement ends, by name?
  • Can it be tested independently of the deployment it came from?
  • Can it be versioned, deprecated and removed?
  • Can it live outside the core behind an extension contract instead?
  • Does moving it into the core reduce total delivery cost across the portfolio, rather than only in the next engagement?

The customization debt register

Every customization that is not immediately promoted or retired goes into a register. This is the artifact that makes upgrade planning possible, and its absence is why organizations discover during a major version upgrade that they cannot enumerate what they have built.

Fields of a customization debt register
FieldWhy it is there
CustomizationWhat was built and where the code lives. A repository link, not a description.
Customer dependencyWhich deployments would break if it were removed. Determines whether removal is a decision or a negotiation.
OwnerA named engineer and a named team. A team alias is not an owner.
Reusable potentialAssessed against the nine promotion questions, not by enthusiasm.
Core API dependencyWhich platform interfaces and which versions. This is the field that makes blast-radius analysis possible before an upgrade.
Upgrade riskWhat breaks at the next major version, assessed rather than assumed.
Maintenance effortObserved, not estimated. Effort that nobody measures gets attributed to delivery and disappears.
Target dispositionPromote, keep as an extension, replace with product configuration, or retire — with a date. An entry with no disposition stays forever, because nothing ever forces the decision.

10 When the product is AI

Forward deployment of a deterministic application is a hard integration problem. Forward deployment of an AI system is that plus a component whose behaviour is part of the production system and changes when the vendor ships a model. Everything below exists because of that difference.

Two claims are worth stating before the diagram, because architectures in this space are frequently built on their opposites. Model capability is not system reliability: a model that answers well in a demo is evidence about the demo. And benchmark performance does not predict workflow performance — a randomised controlled trial published by METR in 2025 found experienced open-source developers took 19% longer to complete real issues when allowed to use AI tooling, while predicting a 24% speedup beforehand and still believing they had been 20% faster afterwards.[69] That result is a snapshot of early-2025 tooling on a specific kind of work and should not be generalized further than that. What generalizes is the shape of the error: the people closest to the system were confidently wrong about its effect, in the same direction, and only measurement caught it.

Select a stage of the production path for the engineering concerns that attach to it.

Prompt injection is a design constraint, not a pending fix

An AI system that reads untrusted content and can take actions has an attacker-controlled input path into its instructions. This is not a solved problem and no credible source claims it is. OWASP’s LLM Top 10 states it directly: “given the stochastic influence at the heart of the way models work, it is unclear if there are fool-proof methods of prevention for prompt injection”.[42] OpenAI’s Chief Information Security Officer described it in October 2025 as “a frontier, unsolved security problem”.[47] Anthropic, publishing a defence that reduced attack success to 1% against an adaptive attacker, wrote that this “still represents meaningful risk” and that they were sharing the results “to demonstrate progress, not to claim the problem is solved”.[48] The attack class was formally characterized in 2023 and has not been closed since.[49] NIST’s adversarial machine learning taxonomy treats it as an open problem area and warns that “many of these mitigations may themselves be vulnerable to new discoveries and evolutions in attacker techniques”.[30]

The architectural consequence is specific. Split your controls into two sets and be honest about which is which. Input filtering, provenance framing and content classification reduce the rate at which injection succeeds; they are classifiers operating on adversarial natural language and cannot carry a safety case. Authorization enforced in the downstream system, identity and audience binding, egress control and approval bound to a specific payload do not care whether the model was fooled. Put the safety case on the second set and budget the first as defence in depth.

Tool integration and MCP

Where a system reaches tools through the Model Context Protocol or an equivalent, the protocol specification itself is unusually clear about what it does not do. Revision 2026-07-28 states that “descriptions of tool behavior such as annotations should be considered untrusted, unless obtained from a trusted server”, and that “while MCP itself cannot enforce these security principles at the protocol level, implementors SHOULD” build the controls themselves.[50] It also carries hard normative requirements that early integrations routinely violated, including “MCP servers MUST NOT accept any tokens that were not explicitly issued for the MCP server” and “MCP servers MUST NOT treat possession of a state handle as authentication”.[50]

The risk is not theoretical. CVE-2025-49596 in the MCP Inspector developer tool carried a CVSS v4.0 base score of 9.4 for unauthenticated remote code execution caused by a missing authentication step between the inspector client and its proxy.[51] A forward deployed engagement that connects an agent to customer systems inherits every one of these decisions. The minimum set worth insisting on: an inventory of which servers are trusted and who approved them, per-target token exchange with an explicit audience rather than passthrough, tool permissions scoped to the single call, secrets held outside the tool’s reach, validation of the action rather than of the intent behind it, transaction and rate limits, and an audit record that survives the compromise of the thing being audited.

OWASP published a Top 10 for Agentic Applications in December 2025 with its own identifier series — ASI01 Agent Goal Hijack through ASI10 Rogue Agents — which is a better fit for this material than the LLM list alone.[44] There is also an OWASP MCP Top 10, but it is an Incubator project at version 0.1 and should not be cited with the weight of a flagship document.[45]

11 GraphRAG: useful grounding, not a guarantee

A semantic layer over enterprise data is a good idea for reasons that have nothing to do with hallucination. It is also routinely sold on a claim that the measurements do not support, and the correction matters because the claim changes what controls people think they need.

The claim I am correcting. One of the source documents behind this project states that “to achieve zero hallucinations, the LLM is forced to write a formal graph database query… based only on the schema”, and that integrating an LLM with an ontology “shifts AI architecture from probabilistic guessing to deterministic reasoning”. Neither is true, and the second document in the same family contradicts the first by describing a verification agent whose existence presupposes that the first stage produces errors.

Notably, the vendor most associated with this architecture does not make the claim either. Palantir’s own engineering post on the subject is titled Reducing Hallucinations with the Ontology, describes the ontology as helping to ground responses, and pairs it with human oversight.[53] The peer-reviewed literature uses the same verb: a NAACL 2024 survey of knowledge graphs and LLMs frames the contribution as mitigating hallucination, not removing it.[67]

What a graph or ontology genuinely provides

  • Retrieval constrained to entities and relationships that exist, which removes a class of fabricated references.
  • Provenance that points at identified nodes and edges rather than at a text chunk, which is the strongest governance argument for the pattern.
  • A structured place to enforce authorization, including permissions inherited through relationships rather than restated per object.
  • Better handling of relationship-heavy questions than passage retrieval, which is what graph structures are actually for.
  • Normalization across heterogeneous sources, with the trade-off stated: normalizing an EC2 instance and an Azure VM into one concept discards the provider-specific attributes that some questions need.
  • An execution step that is genuinely deterministic. Given a fixed query and a fixed graph state, the result is reproducible and auditable.

What it does not provide

The pipeline has three stages and only the middle one is deterministic. Translating a question into a query is a language-model operation. Turning the returned rows back into prose is a language-model operation. Both are probabilistic, and constraining the output to a valid schema does not change that — it changes the failure mode from fabricated prose to a query that is syntactically valid, schema-conformant, and answers a different question than the one asked, returning a clean-looking result set with no signal that anything went wrong.

The confident-wrong-answer failure path A natural-language question is interpreted incorrectly, producing a syntactically valid but semantically wrong graph query, which executes successfully and returns a valid graph result, which is rendered as a confidently incorrect answer. Two markers indicate that schema validation passes at the query stage and that execution succeeds, so neither control detects the fault. Natural-languagequestion Incorrectinterpretation Valid query,wrong question Valid result set Confidentlyincorrect answer schema validation passes here execution succeeds here Neither of the two automatic checks in this pipeline can see the fault. Only comparison against expected output can.
Figure 5 — The failure path that schema conformance does not catch. Both intermediate checks report success.

The numbers

These are the measurements that made me rewrite this section rather than soften it.

Measured accuracy of query generation and related tasks
MeasurementResultSource
Text-to-Cypher execution accuracy76.8% overall for the strongest model tested; 49.9% on complex aggregationPurpose-built Cypher benchmark, EMNLP 2025 Industry Track[56]
Schema validity against execution accuracy0.916 schema-valid, 0.189 correct answers — same model, same tasksEnterprise-schema evaluation, 2026 preprint[57]
Text-to-SQL, the mature analogueBest systems around 82% execution accuracy against a human baseline of 92.96%BIRD benchmark leaderboard[59]
Errors that execute cleanly36.1% of errors on one benchmark are semantic faults that run without error and return wrong dataError taxonomy, PACM SE / FSE[58]
Entity resolution on unseen entities89.0 F1 on entities seen in training, 64.6 F1 on unseen — roughly a 25-point dropWDC Products benchmark, EDBT 2024[66]
Hallucination with retrieval present43.1% of responses in one RAG corpus contained at least one hallucinationRAGTruth, ACL 2024[60]
Groundedness when the source is in the promptBest systems in the low-to-mid 80s%; nothing reaches 100%FACTS Grounding, Google DeepMind — vendor-reported[61]

The 0.916 against 0.189 pair is the one to keep. A schema checker passing more than nine in ten generated queries while fewer than one in five return the right answer is the arithmetic of why schema verification is not error elimination. It catches references to labels and properties that do not exist, which are the cheapest errors. The expensive ones use the right vocabulary to ask the wrong question.

Entity resolution deserves separate attention because a knowledge graph inherits every merge error and every missed merge as a fact. The benchmark drop from 89.0 to 64.6 F1 falls precisely on the case an enterprise graph meets every time a new supplier, customer or asset appears.[66] That is not an argument against building the graph. It is an argument for treating it as a maintained asset with a measured resolution quality rather than as ground truth.

Nor does an LLM verification step close the gap on its own. Google DeepMind’s ICLR 2024 result is that language models “struggle to self-correct their responses without external feedback, and at times, their performance even degrades after self-correction”.[65] A schema is external feedback about schema errors, which is why that check works. It is not external feedback about whether the query captured the user’s intent.

Mitigations that actually apply

01 Schema-constrained generation

Present a pruned, relevant schema rather than the whole model. Neo4j’s own research found that reducing schema size improved accuracy for most models tested and reduced cost for all of them, though not uniformly and from a low base.[68]

02 Query validation before execution

Reject references to labels, properties and relationships that do not exist. Cheap, worth having, and explicitly insufficient on its own.

03 Access-aware retrieval

Evaluate the requesting principal’s permissions at the enforcement point, not by trusting a claim the model passed along. Retrieval that ignores authorization turns a search feature into a data exfiltration path.

04 Deterministic rules where they exist

If a business rule can be expressed in code, express it in code. Asking a model to apply a rule you could have written is choosing a probabilistic implementation of a deterministic requirement.

05 Result validation

Check the returned set against expectations that do not come from the model — cardinality, type, range, referential consistency. This is the only automatic check positioned to catch the right-vocabulary-wrong-question failure.

06 Abstention as a first-class outcome

OpenAI’s position is that hallucination persists partly because “standard training and evaluation procedures reward guessing over acknowledging uncertainty”.[64] Build a path for “I cannot answer this from the available data” and evaluate it as a success.

07 Human review for consequential actions

Rationed to actions that are irreversible or high-value, because approval degrades under volume and an approval queue that is always full stops being read.

08 Continuous evaluation

Retrieval quality and generation quality measured separately. ICLR 2025 work found models fail differently depending on whether the retrieved context was sufficient, and that insufficient context is worse than none — one model’s incorrect-answer rate rose from 10.2% with no context to 66.1% with insufficient context.[62]

On GraphRAG specifically. Microsoft’s GraphRAG is worth reading for what its authors claim, which is narrower than how it is usually cited. The paper’s headline result is improvement in comprehensiveness and diversity of answers to global sensemaking questions over a corpus, evaluated by preference judgement, not factual accuracy.[54] The project’s own responsible-AI documentation states that “human analysis by a domain expert of the answers is needed in order to verify and augment GraphRAG’s generated responses” and that “the system is designed for trusted users”.[55] Hallucination there is measured as a rate.

12 Zero Trust is broader than data access

Access control at the semantic layer is a real Zero Trust control and a good one. It is not Zero Trust, and the difference is not pedantry — it decides which controls a design review will find missing.

NIST SP 800-207 defines zero trust architecture as “an enterprise’s cybersecurity plan that utilizes zero trust concepts and encompasses component relationships, workflow planning, and access policies”, and states plainly that “ZT is not a single architecture but a set of guiding principles” and that transitioning to it “cannot simply be accomplished with a wholesale replacement of technology”.[21] The document lists seven tenets. Role- and attribute-based authorization enforced at a data or ontology layer contributes strongly to three of them — per-session access decisions, dynamic policy incorporating identity and asset state, and authorization strictly enforced before access is allowed. It contributes nothing on its own to the other four: treating all data sources and computing services as resources, securing all communication regardless of network location, continuously measuring the integrity and security posture of every asset, and collecting state telemetry to improve posture.

The Zero Trust control stack Ten control domains arranged in two columns. The left column, from top: human identity, workload identity, device and runtime trust, authorization, data policy. The right column: network and service policy, secrets, software supply chain, observability, continuous re-authorization. Three domains — authorization, data policy and continuous re-authorization — are marked "partly"; the remaining seven are marked "no" and require separate mechanisms. CONTROL DOMAIN ADDRESSED BY DATA-LAYER AUTHORIZATION ALONE? Human identity Workload identity Device and runtime trust Authorization Data policy Network and service policy Secrets Software supply chain Observability Continuous re-authorization partly partly partly nonono nononono Green outline plus the word “partly” marks the three domains a semantic-layer policy engine partly serves. Seven need separate mechanisms, separate owners and separate evidence. Meaning is carried by the label, not by the colour.
Figure 6 — Ten control domains, which is a different cut from NIST’s seven tenets. Continuous re-authorization of a request is partly served here; continuous measurement of asset posture belongs to device and runtime trust and is not. A policy engine at the data layer is one enforcement point inside a Zero Trust architecture, and a good one. It is not the architecture.

The practical consequence for a forward deployed engagement is that the security conversation cannot be delegated to whoever owns the data model. Workload identity, secrets custody, supply chain provenance and the telemetry that makes continuous verification possible all sit with different teams, usually on the customer side, and all of them need to be in the architecture before the first production deployment rather than after the first assessment.

13 The access boundary

This is the diagram I would put in front of a customer security team first, because it answers the question they are actually asking: what can your engineer do in my production environment, and how would I know?

Forward deployed access boundary Four zones. On the left, the customer trust zone containing the customer identity provider, customer data, production systems and the customer audit store. In the centre, an access broker containing approved device, federated identity, multi-factor authentication, privileged access management with just-in-time elevation, a scoped role, time-limited authorization and session logging. To the right of the broker, four deployment environments: development, test, staging and production. On the far right, the vendor zone containing source control, signed artifacts, an update service and support telemetry where permitted. Production is outlined in red to mark it as the environment no standing vendor credential should reach. Arrows show that all engineer access passes through the broker, that only signed artifacts move from the vendor zone into deployment environments, and that telemetry leaves only through a controlled egress path. CUSTOMER TRUST ZONE Customer identity provider Customer data Production systems Customer audit store ACCESS BROKER • Approved, managed device • Federated identity + MFA • Request through customer workflow • Privileged access management • Just-in-time elevation • Scoped role, least privilege • Time-limited authorization • Session recording • Independent audit custody No standing production access DEPLOYMENT ENVIRONMENTS Development Test Staging Production Separate credentials per environment. No shared roles. VENDOR ZONE Source control Signed artifacts Update service Support telemetry Telemetry only where contractually permitted. 1. Every engineer path into a customer environment crosses the broker. There is no second route, including for incidents. 2. Only signed, provenance-attested artifacts cross from the vendor zone into any environment. 3. Telemetry leaves through a controlled egress path, subject to contract, and never carries customer data by default.
Figure 7 — The access boundary. Standing administrator access is not the default operating model for this role, and an architecture that assumes it will not survive a serious third-party review.

Two things make this diagram hold up in practice rather than only in a document. Audit custody sits outside the administrative domain of the people being audited, so the evidence survives the case where the vendor account is the thing that was compromised. And access is requested through the customer’s own workflow rather than the vendor’s, which is the difference between the customer being able to revoke access and the customer being told that access was revoked.

Everest Group makes the sharper version of the point, and it is worth taking seriously precisely because it comes from outside the enthusiasm: introducing forward deployed engineers into an environment with mature change control “can bypass these safeguards”, and “real-time changes to production code create security risks and the potential for significant failures”.[14] The correct response is not to argue with that. It is to design the engagement so that the claim is false in your case, and to be able to show how.

14 Where the model fails

Twelve failure modes. Each has symptoms you can look for, a root cause that is usually structural rather than individual, and an owner — because a mitigation with no owner is a hope.

Snowflake deployments Structural

Symptoms. Upgrades are scheduled per customer. Nobody can say what version any given deployment is on without asking. A platform release note has to be rewritten for each account.

Root cause. No extension contract, so every requirement was met by modifying whatever was nearest. The core and the customization were never separated because separating them cost time in week three.

Engineering consequence. The product becomes N products. Upgrade cost grows with the number of customers rather than staying flat, which is the specific economic property a software business depends on.

Mitigation. A versioned extension surface; a rule that custom code cannot reach into core internals; continuous integration for every extension against the next core release, with deployment blocked on failure; a customization debt register with a disposition per entry.

Owner. Platform engineering owns the contract. Delivery leadership owns whether the rule is enforced under pressure.

Hero engineering Structural

Symptoms. One engineer is on every incident call for an account. Documentation is thin because the person who would write it is the person answering the questions. A week of leave becomes a risk event.

Root cause. Ownership was assigned to a person instead of a team, and the incentive structure rewarded responsiveness over transferability.

Engineering consequence. The deployment becomes unmaintainable by anyone else, which means it also becomes unsellable as a repeatable offering.

Mitigation. Pods with shared ownership and deliberate rotation. Runbooks written during the engagement rather than at the end. A handoff exercise where someone else operates the system for a period while the original engineer stays silent.

Owner. Delivery management, not the engineer.

Permanent privilege Security

Symptoms. A vendor account with production administrator rights and no expiry. Access reviews that return “still required” every cycle without evidence. Incident response that depends on that account existing.

Root cause. Implementation access was granted broadly to unblock delivery and nobody was accountable for narrowing it once delivery finished.

Engineering consequence. A vendor credential compromise becomes a customer production incident, and the customer’s own privileged access model has an exception in it that its assessors will find.

Mitigation. Just-in-time elevation through the customer’s workflow, scoped and time-boxed, with session recording and audit custody outside the vendor’s reach. Access should narrow as the deployment matures; if it has not narrowed in a year, that is a finding.

Owner. Customer security, with the vendor obliged to design for it from the first sprint.

Prototype permanence Delivery

Symptoms. The demo is in production. Credentials in configuration files. No tests around the part everyone depends on. A service whose name still contains “poc”.

Root cause. The prototype worked, the sponsor saw it work, and no gate existed between working and deployed.

Engineering consequence. Every subsequent change is high-risk because nobody knows what the system does under conditions the prototype never met.

Mitigation. Treat production readiness as a distinct gate with an explicit checklist and a named approver. See Production readiness. Build prototypes so they are cheap to throw away, and say out loud at the start that they will be.

Owner. Whoever signs the production readiness review. If nobody signs it, it is not a gate.

Scope gravity Delivery

Symptoms. The engineer is fixing problems in adjacent systems they did not build. Requests arrive directly rather than through any intake. The original success measure has not been discussed for two months.

Root cause. An engineer who is present, capable and already inside the estate attracts work. Nothing about this is the customer behaving badly.

Engineering consequence. The engagement outcome is never reached, because effort went to a series of smaller outcomes nobody is measuring.

Mitigation. Written scope with explicit exclusions. An intake path for new requests. A standing agenda item comparing effort against the agreed outcome. Willingness to decline, backed by delivery management rather than left to the individual.

Owner. Delivery management and the customer sponsor jointly.

Product feedback black hole Structural

Symptoms. The same workaround appears in three deployments. Field engineers stop filing issues because nothing happens to them. Product roadmaps contain nothing traceable to a deployment.

Root cause. No route from field observation to roadmap that has a named owner and a response commitment on the product side.

Engineering consequence. This is the failure that turns the whole model into consulting. Marginal deployment cost never falls because nothing learned is ever built once.

Mitigation. A field findings log with a triage commitment. A metric on the product side counting roadmap items that originated in the field. Rotation of engineers between field and product in both directions.

Owner. Product leadership. Delivery can file findings; only product can act on them.

Customization debt Structural

Symptoms. Local requirements accumulate faster than reusable abstractions. A growing share of engineering time goes to maintaining things built for one customer. Nobody can produce a list of what exists.

Root cause. Customization is easy to create and nobody is accountable for retiring it. There is no register, so there is no visibility, so there is no decision.

Engineering consequence. Capacity is consumed by maintenance that was never planned or funded, and it is invisible in delivery reporting because it is booked as delivery.

Mitigation. The register described in Productization economics, reviewed on a cadence, with an explicit retirement decision per entry and time allocated to act on it.

Owner. Platform engineering, with delivery contributing entries.

IP boundary failure Commercial

Symptoms. Nobody can say cleanly which code is the vendor’s reusable intellectual property and which encodes the customer’s process. Contract renewal surfaces a question nobody has an answer to.

Root cause. Repository structure, licensing and contract language were not designed to keep the two separable, and the separation is much harder to establish retroactively than to maintain.

Engineering consequence. Reuse becomes legally risky in one direction and customer confidentiality becomes risky in the other. Both chill exactly the behaviour the model depends on.

Mitigation. Separate repositories with an explicit interface between them. Contract language covering ownership of work product, derived improvements and the vendor’s right to generalize a pattern. Review at engagement start, not at renewal.

Owner. Legal and product jointly, informed by architecture.

Compliance by architecture Governance

Symptoms. A slide claiming the design is compliant with a named regime. Technical controls presented as evidence with no assessment behind them. Nobody can produce the control narrative an assessor would ask for.

Root cause. Confusion between implementing a control, operating it, evidencing it, and having it assessed. These are four different things done by four different sets of people.

Engineering consequence. A gap discovered during assessment rather than during design, at the point where remediation is most expensive and most public.

Mitigation. The vocabulary in Architecture versus compliance. Architecture supports controls; organizations operate them; evidence demonstrates them; assessors assess them; authorizing officials accept residual risk.

Owner. Governance, risk and compliance, with architecture supplying mechanisms and evidence.

AI without evaluation AI

Symptoms. Quality is assessed by demonstration. No labelled evaluation set exists. A model version change ships without a regression run because there is nothing to run.

Root cause. Evaluation was treated as testing, and testing was scheduled after the thing worked rather than as the definition of working.

Engineering consequence. Nobody can tell whether the system got worse, so nobody can tell whether a change was safe. The model provider ships an update and the failure surfaces as user complaints.

Mitigation. A frozen evaluation set built from real cases before the prototype is accepted. Separate measurement of retrieval and generation. A regression gate on every model, prompt or retrieval change. Where an LLM is used as a judge, its agreement with human labels is itself measured — it “is not a silver bullet”.[73]

Owner. The engineer who owns the AI component, with the workflow owner supplying the labels.

Agent over-permission AI security

Symptoms. An agent holds a broad token because scoping was fiddly. Tools that can write, delete or transact are available in contexts that only needed to read. Nobody can enumerate the actions the agent can take.

Root cause. OWASP names the three roots directly: excessive functionality, excessive permissions, excessive autonomy.[43] Excessive Agency rose to third in the 2026 edition of its LLM Top 10, from sixth.[43]

Engineering consequence. A successful prompt injection reaches everything the token reaches, and the blast radius was set months earlier by a convenience decision nobody recorded.

Mitigation. Minimum tool set per context. Per-target token exchange with explicit audience rather than passthrough. Authorization enforced downstream, not by the model. Transaction and rate limits. Approval bound cryptographically to the specific payload rather than to the session.

Owner. Security architecture, enforced in the platform rather than in a guideline.

Engineer burnout People

Symptoms. Sustained after-hours incident load concentrated on the same names. Travel that never settles. Attrition that takes deployment knowledge with it.

Root cause. The model structurally concentrates context, customer pressure and operational responsibility in one place, and organizations under-invest in the mechanisms that spread it because those mechanisms slow delivery in the short term.

Engineering consequence. Mission continuity risk. When the person leaves, so does the only complete model of how the deployment works.

Mitigation. Rotation with real handover. On-call that includes people outside the engagement. Capacity planning that treats customer count per engineer as a managed number. Instrumented team-health metrics reviewed in the same meeting as delivery metrics.

Owner. Engineering leadership. This one cannot be delegated downward, because the people best placed to notice it are the people affected by it.

15 Reference scenarios

Composite architecture scenarios. None of the four below describes a client, a deployment or a delivered system. Each is assembled from constraints that recur across an industry, and each exists to show a trade-off rather than a result. There are no named organizations, no counts, no timelines, no savings and no outcomes, because I have none to report and inventing them is how this genre of writing usually goes wrong.

A Healthcare operations

Problem shape. Operational visibility requires combining a clinical record system, scheduling and device telemetry, while access is tightly constrained by role and the data carries the strictest handling obligations in the estate.

Architecture shape. Integration performed locally rather than by exporting records; standards-based clinical interoperability rather than direct database access; an operational semantic model separate from the clinical record; de-identification where the use case genuinely does not need identity; role-aware application surfaces; AI assistance confined to summarizing and surfacing rather than deciding.

The trade-off. Operational freshness against integration complexity and the width of the privacy boundary. Every increase in timeliness pulls more identifiable data closer to more people, and the honest version of this design states which of those it chose and why.

What I will not claim. That such a design is HIPAA compliant. Applicability depends on the system boundary, the covered entity’s own controls, the business associate agreement and an assessment nobody has performed.

B Cross-region financial operations

Problem shape. Analysts need a global operational view while source records remain subject to jurisdictional constraints and internal need-to-know rules that differ by region.

Architecture shape. Regional processing with the query pushed to the data; policy-aware federation returning aggregates rather than records; lineage preserved so an aggregate can be explained; identity-aware drill-down that resolves to detail only for a principal authorized in that jurisdiction.

The trade-off. Central analytical convenience against residency, privacy and internal access boundaries. Federated queries are slower, harder to optimize and harder to debug, and each regional node is a component someone has to patch.

What I will not claim. A regulatory outcome. The source material this project started from asserted an audit passed with zero findings; there is no regulator, scope, period or assessor attached to that claim and it is not repeated here.

C Remote industrial operations

Problem shape. High-volume telemetry is generated where connectivity is intermittent and expensive, and operational decisions cannot wait for a round trip to a cloud region.

Architecture shape. Inference local to the site; buffering that survives an outage measured in days; compressed telemetry prioritized by operational value rather than by timestamp; model deployment as a controlled, signed, versioned artifact; drift detection running locally with retraining requests queued for the next window.

The trade-off. Centralized compute efficiency against local autonomy and deterministic behaviour. The local model will be smaller and less capable than the one you could run centrally, and that is the price of it working when the link is down.

What I will not claim. Prevented failures or financial savings. Prevented incidents are counterfactual by construction; a saving derived from them is an estimate dressed as a measurement.

D Disconnected, high-assurance environment

Problem shape. Software must operate with restricted or absent connectivity, under tightly controlled data movement, where remote vendor access is not available at all.

Architecture shape. A fully local runtime with no dependency on a remote control plane; deployment artifacts signed and verified before load; synchronization through a controlled, reviewed path rather than an open channel; policy decisions made locally because the policy service cannot be reached; observability that collects locally and is exported deliberately.

The trade-off. Centralized management against local sovereignty and assurance. Everything that a control plane would have done for you becomes a local procedure that a local operator has to execute correctly.

What I will not claim. Any classified deployment, agency or network. The source material named a specific classified network in an unattributed narrative; that claim is not repeated, and neither is the security clearance requirement it stated as universal, which is contract- and nation-specific.

16 When should an organization use forward deployment?

Up to nine questions, depending on the path. The honest answer for most problems is one of the six alternatives, and an organization that reaches “forward deployed engineering” for everything has stopped using the question as a filter.

The alternatives are not lesser outcomes. Everest Group’s framing is that the question is “not whether they are good or bad” but “where they fit”, and that in environments with mature change control forward deployed engineers “are not only unnecessary… but they can be actively harmful”.[14] Choosing partner delivery for a repeatable implementation is a better decision than assigning the scarcest engineers you have to work a certified third party could do without them.

17 Operating-model maturity

Six levels. Most organizations running this model are at level 0 or 1 and describe themselves as being at level 3, which is easy to test: the artifacts of the higher level do not exist and can be asked for.

0HEROIC
Individual engineers, undocumented access, manual deployment

What it looks like. Deployment knowledge lives in people. Access was granted informally and has not been reviewed. Releases happen when someone runs something from a laptop. Customer-specific code exists in places nobody has catalogued.

Primary risk. Dependency on people. A departure takes the deployment’s only complete mental model with it.

Next action. Get everything into version control, write down what access exists, and produce one runbook that someone else can follow end to end.

1REPEATABLE
Engagement templates, source control, basic pipelines, defined access

What it looks like. Engagements start from a template. Code is in git with review. A pipeline builds and deploys. Access is defined even if it is broader than it should be. Runbooks exist and are occasionally wrong.

Primary goal. Make a successful deployment repeatable rather than remarkable.

Next action. Introduce a production readiness gate with a named approver, and start the customization register before you need it.

2GOVERNED
Formal intake, architecture and security gates, evaluation, SLOs

What it looks like. Work enters through an intake with qualification. Architecture and security review happen before build rather than after. AI components have evaluation sets. Services have service level objectives someone watches. Reusable components exist and are used.

Primary goal. Control risk without destroying delivery speed. This is the level where the trade-off is real and where governance most often overshoots.

Next action. Measure how long the gates take. A gate that adds weeks and catches nothing gets routed around, and the routing around is invisible until something fails.

3PLATFORMIZED
SDKs, golden paths, policy as code, deployment automation, feedback pipeline

What it looks like. There is a supported way to build the common thing and it is faster than the unsupported way. Policy is expressed as code and evaluated in the pipeline. Deployment is automated across topologies. Field findings reach product through a route with a response commitment. An extension catalogue exists with owners.

Primary goal. Reduce the marginal effort of the next deployment.

Next action. Measure marginal effort. If it is not falling, the platform investment is not yet doing what it was funded to do.

4SCALABLE
Partners and customer engineers handle common deployments

What it looks like. Certified partners or the customer’s own engineers deliver the well-understood patterns. The scarce internal engineers are deployed against problems that are genuinely new. Productization is a funded activity with a backlog rather than a side effect. Engagement analytics are standardized enough to compare accounts.

Primary goal. Concentrate scarce talent where it creates new leverage instead of where it is most requested.

Next action. Check that the partner channel is delivering the same quality. This is the level where quality quietly diverges and nobody measures it.

5LEARNING
Field signal shapes product priorities; deployment is a source of advantage

What it looks like. Deployment telemetry informs architecture decisions. Repeated field patterns become platform primitives on a predictable cadence. The economics of an engagement improve measurably from one to the next, and someone can show the series.

Primary goal. Make deployment itself a source of product advantage rather than a cost of sale.

Honest caveat. I have seen this described more often than demonstrated. The evidence that an organization is here is a chart of marginal deployment cost over time, and most organizations cannot produce that chart.

18 Who owns what

An example operating model, not a universal one. Every organization I have looked at places at least two of these differently, and the value of writing it down is the argument it starts rather than the grid it produces.

Example responsibility assignment across ten roles and sixteen activities. R is responsible, A is accountable, C is consulted, I is informed.
ActivityFDEPlatform engProduct engResearchVendor securityGRCSolutions archCust. product ownerCust. securityCust. operations
DiscoveryRIIIIICACC
ScopingRCIICCCACC
ArchitectureACCICCRCCI
IntegrationRCIIIIICCA
Application developmentACCIIIICII
Model selectionRCCACCICCI
EvaluationACCCCCIRIC
Security reviewCCIIRCCIAI
Data authorizationCIIICCICAR
Production deploymentRCIIIIICCA
Incident responseCCIICIIICA
SLO ownershipCRCIIIICIA
Change managementRCIICCICCA
AdoptionCIIIIICAIR
ProductizationCRACIICIII
HandoffRCIIICIACC

Three placements in this grid are deliberate and are where most disagreement lands. Production deployment and incident response are accountable to customer operations rather than to the vendor engineer, which is the structural expression of the access-boundary argument: the party accountable for production is the party that controls production. And evaluation is accountable to the engineer who built the AI component but responsible to the customer product owner, because the labels that define correct behaviour are the customer’s knowledge and cannot be delegated to whoever wrote the code.

19 Prototype is not production

A working prototype is evidence that a hypothesis held under the conditions you tested. It is not a production system, and the gap between them is a defined body of work rather than a tidying-up exercise. This is the checklist I would put a named approver behind.

Application 10 checks
  • Automated tests covering the paths the workflow actually depends on, not the paths that were easy to test.
  • Error handling that distinguishes retryable from terminal, and surfaces the difference to the caller.
  • Explicit versioning of the deployed artifact, traceable to a commit.
  • A rollback that has been executed at least once, in a real environment, by someone who was not the author.
  • Configuration separated from code, with environment-specific values held outside the artifact.
  • Startup behaviour defined when a dependency is unavailable — fail fast, degrade, or wait, decided rather than inherited.
  • Resource limits set, and behaviour under limit understood.
  • Idempotency for every operation that can be retried.
  • No credentials, endpoints or customer identifiers in source.
  • A named owner in the repository, not a team alias.
Data 7 checks
  • Lineage from every field the workflow depends on back to its system of record.
  • Freshness expectations stated per source, and monitored against.
  • Quality checks that fail loudly rather than propagating nulls into a decision.
  • Retention and deletion defined per dataset, and implemented rather than documented.
  • Classification applied, with handling rules that follow the classification through transformation.
  • Backfill and replay behaviour defined for the case where a source was wrong.
  • A stated position on what happens to derived data when the source record is deleted.
AI components 8 checks
  • An evaluation set built from real cases, frozen, versioned, and owned by someone who understands the workflow.
  • A recorded baseline for every quality dimension you intend to defend.
  • Regression gates that run on model change, prompt change and retrieval change — all three, independently.
  • Groundedness measured separately from task success, because retrieval failure and generation failure need different fixes.[62]
  • Tool-use tests that assert what the system does not call as well as what it does.
  • Safety and refusal behaviour tested against adversarial inputs, including content that reaches the model through retrieval.
  • An abstention path that is evaluated as a correct outcome rather than a failure.
  • Model version, prompt version and retrieval configuration recorded on every request, or the telemetry is not diagnostic.
Identity 5 checks
  • Distinct service identities per component, with no shared credential across environments.
  • Human roles defined against duties rather than against convenience.
  • Least privilege verified by testing that an over-broad action fails, not by reading the policy.
  • An access lifecycle with joiners, movers, leavers and a review that can result in removal.
  • No standing vendor access to production, and a documented just-in-time path that has been exercised.
Security 7 checks
  • A threat model that names the assets, the boundaries and the assumptions, reviewed with the customer’s security team.
  • Dependency scanning in the pipeline, with a policy for what blocks a release.
  • Secrets held in a managed store with rotation that has been performed rather than scheduled.
  • Encryption in transit and at rest, with key custody separated from platform administration.
  • Security logging that reaches the customer’s own detection capability, in a format it can consume.
  • Integration with the customer’s incident process, including who is called and what they are authorized to do.
  • Signed build artifacts with provenance, and a verification step that actually rejects an unsigned one.
Reliability 6 checks
  • A service level objective that reflects what the workflow needs, agreed with the person who owns the workflow.
  • Capacity understood at peak, not at average.
  • Failure modes enumerated per dependency, with the chosen behaviour for each.
  • Retry strategy with backoff and a bound, so a downstream outage does not become a self-inflicted denial of service.
  • Backup and restore proven by restoring, in a drill, with the time recorded.
  • A disaster recovery position that names what is lost and how much time it takes, rather than asserting that nothing is.
Operations 5 checks
  • Dashboards that answer “is it working” before they answer “what is it doing”.
  • Alerts tied to symptoms users would notice, with a documented action for each.
  • Runbooks written by someone and executed by someone else at least once.
  • An escalation path that reaches a person, with hours and expectations stated.
  • Named operational ownership on the customer side, agreed before go-live rather than discovered during the first incident.
Governance 5 checks
  • An approved-use statement for the system, including what it is explicitly not for.
  • A recorded data-use decision covering purpose, lawful basis where applicable, and any secondary use.
  • An AI risk decision recorded against whatever framework the organization uses, with the residual risk accepted by a named person.
  • Release approval with an accountable approver rather than a group.
  • Evidence retention defined, so the record still exists when it is asked for.
Cost 5 checks
  • Infrastructure cost attributed to the workload rather than to a shared account.
  • Model and token spend measured per successful workflow, not per call.
  • Observability cost budgeted, because high-cardinality AI telemetry is expensive and gets cut first when nobody planned for it.
  • Data movement cost understood, especially in federated and hybrid designs.
  • Ongoing support burden estimated and staffed, or the engagement never ends.
Adoption 5 checks
  • A named workflow owner on the customer side who wants this to exist.
  • Training that reaches the people who will use it, not only the people who sponsored it.
  • A user experience tested with users rather than with the project team.
  • A success metric agreed in advance, measurable without a special report.
  • A feedback mechanism that produces work rather than sentiment.

20 Measuring outcomes without rewarding heroics

No benchmark values appear below. Every organization’s baseline differs and a published target would be a number I invented. What each metric is for is stated instead, because a metric whose purpose is unstated gets optimized in whatever direction is easiest.

Delivery

Time to validated prototype
Reveals whether discovery is converging or circling. A long time here is usually a scoping problem, not an engineering one.
Time to production
Reveals the size of the readiness gap. If it dwarfs time to prototype, the prototype was not built toward production.
Blocker age
Reveals dependency on the customer organization. Old blockers are an escalation signal, not an engineering signal.

Adoption

Active users in the target population
Reveals whether the workflow owner’s people actually adopted it, as opposed to whether accounts were provisioned.
Workflow completion rate
Reveals whether the system finishes the job or hands it back part-done, which is the failure users stop reporting.
Repeat usage
Reveals whether it earned a place in the routine. First use is curiosity. Repeat use is the signal.

Reliability

SLO attainment
Reveals whether the target was set honestly. Permanent 100% attainment usually means the objective is not binding.
Incident rate and MTTR
Reveals operational maturity. Watch the trend rather than the level.
Change failure rate
Reveals whether the readiness gate is doing anything.

AI quality

Evaluation pass rate
Reveals regression on a frozen set. Only meaningful if the set has not been edited to make it pass.
Groundedness
Reveals whether answers are supported by retrieved evidence, measured separately from whether they were useful.
Task success
Reveals workflow effectiveness, which benchmark scores do not predict.[69]
Human escalation rate
Reveals where the system stops being autonomous. A falling rate with flat quality is the good direction; a falling rate with unmeasured quality is a warning.

Economics

Cost per successful workflow
Reveals the real unit economics. Cost per call flatters systems that fail cheaply and retry often.
Model and token cost
Reveals where routing or caching would pay. Routing is an established pattern with published trade-offs, and the savings are highly workload-dependent.[72]
Engineering effort per deployment
Reveals whether the platform investment is working. This is the single number that separates a product from a consultancy.

Productization

Reusable component ratio
Reveals how much of a deployment came from the shelf.
Duplicate customizations
Reveals patterns that should have been promoted and were not. The third occurrence is the signal.
Custom code retired
Reveals whether the debt register is a decision-making tool or a list.
Field patterns converted to platform capability
Reveals whether the loop in Figure 3 is closed.

Product feedback

Actionable field findings
Reveals whether engineers are still filing, which they stop doing when nothing happens.
Roadmap items originating in the field
Reveals whether product is receiving. This is the counterpart metric and both are needed.
Regressions found through deployment
Reveals the value of the field as a test surface, which is real and rarely counted.

Team health

Customer load per engineer
Reveals concentration before it becomes attrition.
After-hours incident load
Reveals whether operational ownership actually transferred at handoff or only formally.
Travel burden
Reveals sustainability. Worth tracking even where embedding is mostly remote.
Ownership concentration
Reveals bus factor per deployment. Should be reviewed in the same meeting as delivery metrics, not in a separate one nobody attends.

21 Standards and governance alignment

This is a crosswalk, not a conformity claim. Nothing in this architecture makes an organization compliant with anything. Applicability depends on the system boundary, the organization’s own controls, contractual obligations, the regulatory context and an assessment that has not been performed. The columns are deliberately named “architecture mechanism”, “operational process” and “example evidence” to keep those three things apart.

Versions below were checked against primary sources on 9 August 2026. Several changed within the preceding twelve months in ways that invalidate widely circulated guidance, and those are marked. Filter to the frameworks that apply to your boundary rather than reading all of them.

Crosswalk between forward deployed engineering concerns and named frameworks, with architecture mechanisms, operational processes, example evidence and caveats
FrameworkConcern it speaks toArchitecture mechanismOperational processExample evidenceCaveat
NIST AI RMF 1.0
AI 100-1, Jan 2023
Whether an AI use case has been governed, mapped, measured and managed rather than merely builtModel inventory; use-case classification; evaluation harness; human oversight points; rollback to a prior model or a non-AI pathIntake and risk classification; recorded acceptance of residual risk by a named person; periodic re-review on model changeRisk decision record; evaluation results per release; oversight design rationaleFinal, but NIST states the framework is being revised, with no public draft as of August 2026.[27] The revision was directed by US federal AI policy in July 2025.[28] Do not cite a revision that has not been issued.
NIST AI 600-1
GenAI Profile, Jul 2024
Generative-specific risks including confabulation, data leakage and information integrityGrounding and provenance; output validation; egress control on generated content; abstention pathPre-deployment red teaming; content-integrity review for consequential decisionsAdversarial test results; groundedness measurements; incident recordsNIST uses “confabulation” rather than “hallucination” and flags the risk as sharpest in consequential decision-making.[29] A companion profile to the AI RMF, so its standing follows the revision.
NIST CSF 2.0
CSWP 29, Feb 2024
Whether the engagement has an owner, and whether outcomes are governed rather than assumedThe whole architecture read as Govern, Identify, Protect, Detect, Respond and Recover surfaces — Govern is the function added in 2.0 and the one an engagement is most likely to leave unassignedGovernance function assignment; third-party risk treatment for the vendor engineering teamCurrent and target profiles; supplier risk assessmentCSF 2.0 says explicitly that it “does not prescribe how outcomes should be achieved”.[23] It is a structure for the conversation, not a control set.
NIST SP 800-53
Rev. 5, Release 5.2.0
The control vocabulary an assessor will use for access, audit, configuration and supply chainAC, AU, CM, IA, SC and SR family implementations across the stackControl selection against a baseline; implementation; assessment; continuous monitoringControl implementation statements; assessment results; monitoring outputCite the release, not just the revision. Release 5.2.0 (August 2025) added SA-15(13), SA-24 and SI-02(07) in response to EO 14306.[24]
NIST SP 800-207
Zero Trust, Aug 2020
Per-request access decisions where the network is assumed compromisedPolicy decision and enforcement points; workload identity; per-session authorization; continuous verificationTrust algorithm definition; policy authoring separated from policy administrationPolicy definitions; authorization decision logs; asset posture data“ZT is not a single architecture but a set of guiding principles.”[21] NCCoE SP 1800-35 became final in June 2025 and consolidated the earlier draft volumes.[22]
ISO/IEC 27001
2022, incl. Amd 1:2024
Whether the provider’s own delivery process, including field engineering, sits inside a managed systemTechnical controls in Annex A relevant to access, cryptography, operations and supplier relationshipsISMS scope covering forward deployed operations; risk treatment; internal audit; management reviewStatement of Applicability; risk treatment plan; audit recordsCite as ISO/IEC 27001:2022 including Amendment 1:2024. The transition deadline passed in October 2025, so a certificate against the 2013 edition is no longer valid.[31]
ISO/IEC 42001
AI management systems, 2023
Whether AI is managed as a system with objectives, roles and improvement rather than per projectAI system inventory; impact assessment inputs; lifecycle controlsAI policy; objectives; competence; operational planning; internal auditAI management system documentation; impact assessments; audit recordsISO/IEC 42006:2025 sets requirements for bodies certifying against 42001, which is what makes accredited certification operational.[32] Certification is of the management system, not of a model or an architecture.
OWASP Top 10 for LLM Applications
2026 edition — renumbered
The application-level risks an AI deployment is most likely to carry into productionDownstream authorization; output handling; retrieval permission checks; consumption limitsThreat modelling per use case; security testing including retrieval-borne injectionTest results; findings tracked to closureChanged in August 2026. Identifiers moved from LLM##:2025 to LLM##:2026; Excessive Agency rose to third; a new entry covers hidden context exposure.[43] Any matrix citing 2025 identifiers is stale.
OWASP Top 10 for Agentic Applications
ASI01–ASI10, Dec 2025
Risks specific to systems that plan, remember and act — goal hijack, tool misuse, privilege abuse, cascading failureTool permission scoping; memory and context integrity; inter-agent authentication; action validation and limitsAgent registration and approval; permission review; kill-switch procedureTool permission matrix; action audit trail; incident recordsA distinct series from the LLM list, with its own identifiers. It did not exist a year ago.[44]
MITRE ATLAS
data release v2026.05
Adversary tactics and techniques against AI systems, modelled as ATT&CK is for enterpriseDetection coverage mapped to technique identifiers; telemetry sufficient to see themThreat-informed defence; detection engineering; purple-team exercisesTechnique coverage map; detection test resultsContinuously updated rather than versioned like a standard; techniques are now tagged by platform including agentic AI.[46] Cite the data release and the date.
CIS Critical Security Controls
v8.1, Jun 2024
A prioritized baseline for the operational hygiene an engagement inheritsAsset and software inventory; secure configuration; account and access management; audit log managementImplementation group selection; safeguard implementation and reviewInventory records; configuration baselines; log management evidencev8.1 added a governance function and realigned mappings to CSF 2.0.[34] Stable since June 2024, which not every row here can say.
SLSA
v1.2 — v1.1 retired
Whether the artifact deployed into a customer environment is the one that was built from the reviewed sourceProvenance generation; artifact signing; verification at deployment that rejects unsigned or unattested buildsBuild platform hardening; provenance policy; exception handlingProvenance attestations; verification logs; policy definitionv1.1 is explicitly retired. v1.2 was approved in November 2025 and adds the source track alongside the build track.[35]
OpenSSF OSPS Baseline
release 2026-02-19
Baseline security practice for the open-source components an engagement pulls inDependency inventory and SBOM generation; component provenance; vulnerability scanning in the pipelineComponent intake policy; upgrade cadence; end-of-life handlingSBOMs; scan results; component approval recordsDate-versioned rather than semantically versioned; the February 2026 release added controls and expanded external mappings.[36]
SOC 2
TSC 2017, points of focus rev. 2022
How a customer gains assurance over controls at a vendor that has access to its systemsLogical access controls; change management; monitoring; the access-broker design in Figure 7Control operation over a defined period; management’s description and assertionThe examination report itself; control operation evidence for the periodSOC 2 is an examination, not a certification. There is no such thing as being “SOC 2 certified”, and it says nothing about whether an architecture is sound.[33] Security is the mandatory category.
HIPAA Security Rule
45 CFR 164 Subpart C
Only where electronic protected health information is inside the boundaryAccess control, audit controls, integrity, transmission security; de-identification where the use case permitsRisk analysis; workforce and business associate management; sanction and review policiesRisk analysis record; BAA; access review evidenceThe proposed 2025 overhaul has not been finalized and has moved to the long-term actions list; the existing Security Rule governs.[37] Guidance predicting a 2026 final rule is now wrong.
PCI DSS
v4.0.1
Only where cardholder data enters the system boundarySegmentation that removes systems from scope; tokenization; MFA for all access to the cardholder data environmentScope confirmation; targeted risk analyses; authenticated internal scanningScope documentation; scan results; the assessment itselfThe previously future-dated v4.x requirements have been in force since 31 March 2025.[38] Describing them as upcoming is a year out of date.
FedRAMP
20x, Phase 3 active
Only for cloud offerings serving US federal agenciesCloud-native architecture, identity, logging and recovery expressed as continuously validated indicatorsContinuous validation rather than point-in-time authorization packagesMachine-readable evidence against Key Security IndicatorsTerminology changed. “Authorization” is now Certification, impact levels are Certification Classes A–D, the SSP is a Security Decision Record and POA&M items are Accepted Weaknesses.[39] There is no mechanism by which an architecture is pre-accredited.
CMMC and NIST SP 800-171
Phase II suspended Jul 2026
Only where controlled unclassified information is in scope for a defence supply chainThe 800-171 requirement families as implemented in the deploymentSelf-assessment, scoring and affirmation under the phase currently in forceAssessment score; system security plan; affirmation recordCMMC Phase II was suspended on 13 July 2026; only Phase 1 self-assessment is in force.[40] Note also that NIST’s current publication is SP 800-171 Rev. 3 while CMMC Level 2 is still assessed against Rev. 2.[41]

22 Architecture versus compliance, worked once

The distinction the previous section rests on is easier to see in a single example than in an argument. Privileged access is the right one to use, because it is the control a forward deployed engagement is most likely to weaken and the one an assessor will look at first.

Worked example separating architecture mechanism, operational process, evidence and assurance for privileged access
LayerPrivileged access, worked through
Requirement areaVendor engineers need occasional access to a customer production environment to diagnose and repair.
Architecture mechanismFederated identity from the customer’s provider with multi-factor authentication; a just-in-time role that does not exist until requested; privileged access management brokering the session; session recording; audit records written to a store the vendor cannot administer.
Operational processAn access request raised in the customer’s own workflow with a stated reason and duration; approval by someone other than the requester; automatic expiry; periodic review that can and sometimes does result in removal.
EvidenceThe access request and its approval; the authorization decision log; the role definition showing scope; the session recording; the access review record showing what was removed.
AssuranceSomeone independent tests whether the control operated as designed over a period — that expiry actually expired, that reviews actually removed access, that the recording exists for a sampled session.
AuthorizationAn accountable person accepts the residual risk. That decision is theirs and cannot be produced by a diagram.

Each row is a different activity performed by different people. The architecture supplies the second row and makes the fourth row possible. It cannot supply the third, the fifth or the sixth. NIST puts the responsibility explicitly on the organization: “organizations have the responsibility to select the appropriate security and privacy controls, to implement the controls correctly, and to demonstrate the effectiveness of the controls in satisfying security and privacy requirements”.[24] Assessment is defined as determining “if the controls are implemented correctly, operating as intended, and producing the desired outcomes”.[25] None of those verbs belongs to a design.

“Security and privacy control assessments are not about checklists, simple pass/fail results, or generating paperwork.”

NIST SP 800-53A Rev. 5[26]

23 What a disciplined engagement produces

Nineteen artifacts. None of them are downloadable from this page, because I have not written them for a real engagement and publishing templates I have not used would be filler. What follows is what each one is for and why its absence hurts.

01 Engagement charter

States the outcome, the boundary, the exclusions and who decides. The document people stop reading in month two and should reread in month four.

02 Discovery findings

What was observed rather than what was requested. The gap between the two is usually the whole value of the discovery phase.

03 Context diagram

The system, the actors and the neighbouring systems on one page. If it needs two pages, the boundary is wrong.

04 Architecture decision records

Decision, context, options considered, consequence. Their value is in year two, when someone asks why and the person who knew has left.

05 Integration inventory

Every system touched, its owner, its protocol, its availability and its data classification. The artifact that makes change impact assessable.

06 Data-flow diagram

Where data moves, crosses a boundary, is transformed and is retained. Prerequisite for both the threat model and the privacy analysis.

07 Threat model

Assets, boundaries, adversaries, assumptions. Its most useful section is the assumptions, because those are what change.

08 Identity and access matrix

Who and what can do what, in which environment. The document that makes over-permission visible instead of theoretical.

09 Data classification map

Classification per dataset and the handling that follows from it, tracked through transformations rather than stated at the source.

10 Evaluation plan

For AI components: what is measured, on which frozen set, by whom, and what result blocks a release.

11 Production readiness review

The checklist in section 19 with a name against it. If nobody signs it, it has not gated anything.

12 SLO definition

What the workflow needs, agreed with the person who owns the workflow, with the error budget and what happens when it is spent.

13 Risk register

Open risks with owners and treatment. Distinct from the issue log; a risk that has occurred is an issue and should move.

14 Cutover plan

Sequence, dependencies, verification steps, rollback trigger and the person authorized to pull it. Written before the day.

15 Operations runbook

Written by one person and executed by another before go-live. A runbook nobody has followed has not been tested.

16 Incident escalation matrix

Who is called, in what order, with what authority, and what the vendor is permitted to do without waiting.

17 Customization debt register

The register from section 09. The artifact that makes upgrade planning a calculation rather than an archaeology exercise.

18 Product feedback log

Field findings with a triage commitment attached. Without the commitment, the log stops receiving entries, and the engineers who stopped filing were right to.

19 Handoff package

Everything above, plus the demonstration that someone else operated the system. Handoff has to be proven by someone else running the system, which is why it belongs in this list rather than in a folder.

24 What the model teaches

Eight conclusions I would defend, written as claims rather than as principles because principles are harder to disagree with and disagreement is the useful part.

  1. Proximity to the customer is not a substitute for engineering discipline, and is often used as one. The argument that tests, reviews and documentation can wait because the customer needs it Friday is the same argument every quarter, and it compounds.
  2. The hardest problems sit between systems and between teams, not inside an API. That is why the role exists at all, and why staffing it with someone who cannot change code produces a well-documented list of things that cannot be done.
  3. A successful prototype is evidence that a hypothesis held, not evidence that a system exists. Treating the two as the same thing is how a project comes to overrun its estimate by a multiple rather than a margin.
  4. The model scales only when repeated fieldwork becomes reusable platform capability. Everything else in the operating model is downstream of whether that loop is closed and instrumented. Nobody has ever closed it by intending to.
  5. Access should narrow as a deployment matures. If the vendor account still holds the permissions it needed during implementation, the engagement has a security posture that was set by delivery pressure two years ago.
  6. AI deployment requires continuous evaluation because model behaviour is part of the production system. A component whose vendor can change its behaviour without your release cycle is a dependency with no change control, and evaluation is the only thing standing in for it.
  7. A semantic layer improves retrieval, provenance and policy enforcement, and does not eliminate probabilistic failure. The measurements are unambiguous and they are worse than the diagrams imply. Design for the wrong-but-plausible answer, because it will occur and it will look correct.
  8. The strongest organizations optimize two things at once: the customer outcome, and the rate at which field learning improves the product. Optimize only the first and you have built an excellent consultancy. Optimize only the second and you have built a platform nobody has deployed.

25 Questions

The role

What exactly is a forward deployed engineer?

A software engineer who works inside a customer’s operational context with the authority to design and build production systems there, and who stays accountable for how those systems behave. The distinguishing variables are depth of technical intervention and persistence of presence, not job title or location.

Is it just another name for consulting?

Sometimes, and the industry does not agree with itself. Traditional consulting typically produces recommendations that a separate team implements; the forward deployed model has the same people design, build and operate. But Databricks files its team under professional services,[8] Vercel’s role reports to a Director of Professional Services while the posting says it is not consulting,[9] and OpenAI and Anthropic both moved delivery into separate services ventures.[6][7] The useful question is not what it is called but whether the engineers have production authority and whether their learning reaches the product.

How is it different from a solutions architect?

A solutions architect is accountable for a design; a forward deployed engineer is accountable for a running system. The architect’s deliverable is a decision; the engineer’s is behaviour in production. Both roles exist for good reasons and the failure mode is asking one to do the other’s job with the other’s authority.

Does the role have to be on-site?

No. Embedding is operational, not geographic. Palantir’s own posting for the role caps travel at 25%,[2] Anthropic’s federal posting states 25–50%,[4] and PostHog runs its forward deployed engineer fully remote and asynchronous by deliberate choice.[12] The one topology where physical presence is genuinely forced is the disconnected environment, and even there the person present may be a cleared local operator rather than the vendor’s engineer.

How much coding is actually involved?

It varies enough that any single figure would be misleading. One independent reading of a major cloud vendor’s posting estimated roughly a quarter coding, half integration work, and a quarter meetings.[13] The number that matters is not the percentage but whether the engineer has authority to merge to production. Without that, the role is advisory regardless of how much code gets written.

What skills actually distinguish an effective one?

Software engineering depth is necessary and not sufficient. The differentiators are the ability to scope an ambiguous problem into something buildable, comfort operating inside another organization’s politics without becoming part of them, willingness to write the boring artifacts, and the judgment to distinguish a requirement that should be built from a request that should be declined.

Can a solutions architect move into the role, or an engineer move back into product?

Both happen and both are useful. Architects moving in need to rebuild the habit of shipping and being on call for what they shipped. Engineers moving to product carry a specific advantage, which is that they have watched the product fail in ways the roadmap did not anticipate. Deliberate rotation in both directions is one of the few structural mitigations for the feedback-loop failure that actually works.

Scope and economics

What makes a problem suitable for this model?

Strategic importance, genuine workflow ambiguity, integration complexity that standard configuration cannot reach, a production outcome that requires software to be written, and a reasonable chance the work reveals a reusable platform capability. Missing the last one is acceptable occasionally and fatal as a pattern.

When should professional services or a partner do it instead?

When the pattern is known and the work is implementation rather than discovery. A repeatable deployment handed to a certified third party is a better use of everyone than assigning your scarcest engineers to it. Everest Group goes further and argues that in environments with mature change control, forward deployed engineers can be “actively harmful” because real-time production changes bypass the safeguards those environments depend on.[14]

Is the model economically viable at scale?

Not as a universal delivery mechanism, and the people who run it say so. The viable shape is selective: forward deployment for landings, frontier problems and reference implementations, with repeatable patterns moving to a lower-cost tier or a partner. F-Prime Capital offers a threshold worth arguing about — if more than 30–40% of deployments need significant forward deployed effort, the problem is product design rather than go-to-market.[15] Treat that as a prompt to measure, not a target.

How do you stop customer-specific code becoming permanent debt?

Register it at creation with an owner, a core API dependency and a target disposition; run continuous integration for every extension against the next core release and block deployment on failure; review the register on a cadence with time allocated to act on it. The mechanism is not complicated. What fails is that nobody is accountable for retirement, so every entry defaults to permanent.

Who owns code produced during an engagement, and how should IP boundaries work?

Decide at engagement start and structure the repositories to match. Vendor reusable intellectual property and customer-specific process logic belong in separate repositories with an explicit interface, and the contract needs to cover work product ownership, derived improvements and the vendor’s right to generalize a pattern. Establishing this retroactively is materially harder than maintaining it, and the moment it is discovered to be missing is usually a renewal negotiation.

Security and access

Should engineers hold production administrator privileges?

Not as a standing arrangement. Just-in-time elevation requested through the customer’s own workflow, scoped to a role, time-boxed, session-recorded, with audit custody outside the vendor’s administrative reach. The cost is minutes at the start of an incident. The benefit is that a vendor credential compromise stops being a customer production incident.

How should access work in regulated environments?

The same way, with more evidence. Co-design the roles with the customer’s security and compliance teams rather than requesting access and negotiating afterwards. Expect to be managed as a third party or sub-processor, expect audit rights in the contract, and expect a joint governance forum reviewing architecture and changes. The engagements that go badly are the ones where security was engaged after the design was fixed.

How does the model align with DevSecOps and Zero Trust?

DevSecOps is the delivery discipline that keeps the model from producing snowflakes: pipeline-enforced testing, scanning, policy as code, signed artifacts and provenance. Zero Trust is the access posture that keeps it from producing standing privilege. Neither is optional and neither is produced by the architecture on its own — see section 12 for what a data-layer policy engine does and does not cover.

How does tool integration through MCP change the security architecture?

It converts an integration layer into something closer to a control plane, with a non-deterministic caller. The specification is explicit that tool descriptions are untrusted unless the server is trusted, and that the protocol cannot enforce its own security principles.[50] Practical minimum: an inventory of trusted servers with named approvers, per-target token exchange with explicit audience rather than passthrough, tool permissions scoped to the call, action validation, transaction limits, and audit that survives compromise of the audited system.

What should happen before an agent takes a real-world action?

Authorization enforced in the downstream system against the acting principal, not a decision made by the model. Validation of the action’s structure rather than its apparent intent. A limit on value and rate. Human approval where the action is irreversible or high-value, bound to that specific payload rather than to the session. And a log entry that records the model version, prompt version, retrieved context, tool call and authorization decision, because without those the incident review has nothing to work with.

AI systems

How is AI forward deployment different from normal application implementation?

Four differences. Quality is statistical rather than binary, so acceptance needs an evaluation set instead of a test suite. Behaviour changes when the model changes, on the vendor’s schedule rather than yours. Cost varies per request. And the system reads untrusted content, which means content is an attack surface. Everything else is a normal, difficult integration project.

Does retrieval eliminate hallucination?

No. One annotated corpus found 43.1% of responses in retrieval-augmented settings contained at least one hallucination.[60] Even when the source document is supplied directly in the prompt, the best measured groundedness scores sit in the low-to-mid eighties.[61] Retrieval substantially reduces the rate. It does not remove the failure mode.

Does a graph or ontology eliminate hallucination?

No, and the vendor most associated with the claim does not make it. Palantir’s own material is titled Reducing Hallucinations and pairs the ontology with human oversight.[53] Peer-reviewed work on knowledge graphs and LLMs frames the contribution as mitigation.[67] There is also a formal argument that hallucination cannot be eliminated in general,[63] and a competing position from OpenAI that confident hallucination is reducible because a model can abstain.[64] Both agree accuracy never reaches 100%.

What role does an enterprise ontology actually play, and is one required?

It gives you shared semantics across fragmented sources, a place to express authorization over relationships, and provenance that points at identified records. Palantir describes its Ontology as an operational layer connecting integrated digital assets to their real-world counterparts.[52] It is not required. It is justified when questions are relationship-heavy, entities are resolvable, and somebody will own the schema. Build one because you need those properties, not because the architecture diagram has a slot for it.

How should evaluation be handled?

A frozen set built from real cases before the prototype is accepted; retrieval and generation measured separately; regression gates on model, prompt and retrieval changes independently; abstention scored as a correct outcome. Where an LLM acts as judge, measure its agreement with human labels first — the practitioner consensus is that it “is not a silver bullet”.[73] Multi-dimensional evaluation is the established academic position, not a novelty.[71]

Why not just trust benchmark scores?

Because they measure benchmarks. A 2025 review found task-setup and reward-design flaws in widely used agentic benchmarks capable of shifting reported performance by up to 100% in relative terms, including one that scored empty responses as successes.[70] And the METR trial found real developers slowed down while believing they had sped up.[69] Evaluate against your own workflow.

Is the system deterministic if the query execution is?

One stage is. Given a fixed query and a fixed data state, execution is reproducible and auditable, which is genuinely valuable. The stages either side of it are language-model inference and are not deterministic — not even at temperature zero in practice, without batch-invariant serving.[74] Claim the determinism you have and no more.

Governance and organization

Does following NIST or ISO make the system compliant?

No. Architecture supports controls; organizations implement and operate them; evidence demonstrates them; assessors assess them; an accountable person accepts residual risk. NIST places selection, correct implementation and demonstration of effectiveness on the organization.[24] A design cannot perform any of those verbs. See section 22.

How do field engineers work with product and research?

Through a route with a named owner and a response commitment on the receiving side. A findings log that product is not obliged to triage stops receiving entries, and the engineers are right to stop filing. Rotation in both directions is the mechanism that keeps the route honest, because it puts people on the receiving side who remember what filing felt like.

What should be productized?

The pattern that has appeared three times, whose abstraction is stable, whose lifecycle has an owner, and whose promotion reduces cost across the portfolio rather than in the next engagement. The step organizations skip is the intermediate one: a supported extension contract where the specific thing can live without becoming core.

How do you prevent burnout?

Treat it as a structural property rather than an individual one, because the model concentrates context, customer pressure and operational load by design. Pods with real shared ownership; rotation with proven handover; on-call that includes people outside the engagement; customer load per engineer managed as a number; team-health metrics reviewed in the same meeting as delivery metrics rather than in a separate one.

What does a mature organization look like?

Marginal deployment effort falls measurably from one engagement to the next, and someone can show the series. Field findings appear in the roadmap with traceability. Partners deliver the known patterns. Access narrows over time rather than accumulating. The customization register has retirements in it. Most organizations that describe themselves this way cannot produce the first artifact on that list.

What are the main risks of adopting the model?

Over-dependence on specific individuals; insufficient governance around access and data use; and failure to route field learning into product strategy. Those three are where the serious damage concentrates, and all three are organizational rather than technical, which is usually why they are addressed last.

26 Corrections

This project began from four research documents on forward deployed engineering and enterprise ontology. Parts of them are good: the deployment-topology analysis, the challenge and mitigation pairs, the lifecycle structure and the governance discussion are all sound and are reflected throughout this page. Parts did not survive verification. Those are listed here rather than quietly dropped, because a reader has no way to tell checked work from written work unless the checking is shown.

Claims in the source material that were corrected, with what was actually found
Claim in the source materialWhat verification found
“To achieve zero hallucinations, the LLM is forced to write a formal graph database query… based only on the schema.”Not supported by any source I could find, and not claimed by the vendor most associated with the pattern. Schema conformance and correctness are measurably different: one evaluation reports 0.916 schema validity against 0.189 execution accuracy for the same model on the same tasks.[57] The phrase is quoted here only as the claim under correction; it is not asserted anywhere on this page.
Integrating an LLM with an ontology “shifts AI architecture from probabilistic guessing to deterministic reasoning”.Overstated. Query execution is deterministic; question-to-query translation and result-to-prose generation are language-model inference and are not, even at temperature zero without batch-invariant serving.[74] Corrected in section 11.
“By integrating the LLM through an ontology, you inherently adopt a Zero Trust architecture.”Refuted. NIST SP 800-207 defines zero trust architecture as an enterprise plan encompassing component relationships, workflow planning and access policies, and states that it “is not a single architecture but a set of guiding principles”.[21] Data-layer authorization addresses roughly three of seven tenets. Corrected in section 12.
“Prevented two catastrophic failures in the first year, saving an estimated $40M.”Unattributed, and methodologically unsound regardless: prevented failures are counterfactual, so a saving derived from them is an estimate presented as a measurement. Removed. The industrial scenario in section 15 carries no outcome claim.
“Achieved full compliance, passed a regulatory audit with zero findings.”No regulator, scope, period or assessor is attached. “Full compliance” is also not a state a system attains. Removed, and replaced by the vocabulary in section 22.
A named classified network, a three-letter agency, a twelve-facility hospital network with named clinical vendors, a North Sea platform with precise bandwidth and sensor counts.Four unattributed narratives carrying precise, quotable numbers. Rewritten as four labelled composite scenarios with every organization, count, timeline and outcome removed. The hospital narrative also endorsed passive interception of clinical message traffic as a workaround for denied API access, which is not a practice I would recommend or repeat.
Security controls and audit trails “automatically generated to satisfy FedRAMP, HIPAA, GDPR, ITAR”; the edge architecture “can be pre-accredited”.No such mechanism exists. Architecture generates evidence; organizations operate controls; assessors assess; authorizing officials accept risk. FedRAMP has also restructured — authorization is now certification, impact levels are certification classes, and the system security plan is a security decision record.[39]
“Secret or Top Secret/SCI clearance is mandatory, often with a polygraph.”Nation- and contract-specific, stated as a property of the role. Not repeated. The disconnected scenario notes only that vendor remote access is unavailable in that topology.
A fixed team topology with an “80% code” split and universal ratios; two adoption timelines differing by a factor of two to three.Both unattributed, and the two source documents contradict each other. This page presents phases and exit criteria without month ranges, and treats team shape as an example rather than a standard.
Forward Deployed Architecture described as “a coherent architectural discipline”.I searched Open Group, IEEE, ISO, SEI, IASA, IETF and NIST material and found no recognition of the term as an architecture discipline. The certifications on offer are commercial training products without standards-body accreditation. “Forward Deployed Architect” does exist as a job title at individual companies. This page therefore avoids the abbreviation entirely and calls the design what it is: a synthesized reference architecture supporting an operating model.
“Kubernetes / K3s / MicroK8s: the universal application orchestration layer”; “interoperability with 100+ legacy protocols”; a policy engine that “prevents exfiltration even by privileged users”.Three technology claims stated as requirements or guarantees. None is defensible as written. The architecture in section 04 names no mandatory technology, and the security section is careful about what a control prevents as opposed to what it detects and constrains.
“In late June 2026, Amazon Web Services announced a $1 billion investment” — appearing only as a truncated reference fragment.Verified. AWS announced it on 30 June 2026 to fund a forward deployed engineering organization embedding engineers with customers for agentic AI work.[1] The claim was correct; only its context was missing. The narrower caveats are that no timeframe is disclosed for the figure, no headcount is given beyond “thousands”, and the accompanying speed claims are vendor-reported and unaudited.
A bibliography attributing specific titles to CNCF, Styra, Microsoft and HHS, and an institutional byline with a classification marking.Several of those citations do not appear to correspond to real publications, and the byline and marking are not real provenance. Nothing from that bibliography is relied on here. Every source in section 27 was retrieved and read.

Two claims I could not settle. The widely repeated figure that postings for the role grew more than 800% between January and September 2025 traces to the Financial Times through secondary coverage, but no source I found identifies the underlying dataset; it also counts postings rather than filled roles from a small base.[19] And Salesforce’s stated commitment to build a team of a thousand forward deployed engineers is a commitment, not a verified headcount.[10] The Salesforce figure appears here and nowhere else. The 800% figure appears once in section 02, carrying the same qualification.

27 Sources and method

This project combines analysis of publicly documented forward deployed operating models, architecture and security standards, peer-reviewed research on retrieval and query generation, and current AI security guidance. The architecture diagrams are reference designs rather than representations of any vendor implementation. The four deployment scenarios are composites and are labelled as such. No client, deployment, customer or measured result is described anywhere.

Method, briefly. I read the four source documents in full and built a claim-by-claim ledger classifying each substantive statement as publicly verifiable, source-derived, an architectural recommendation, a composite scenario, or an interpretation — then verified the first category against primary sources and corrected or removed what failed. Every version number, publication date, framework identifier and quoted normative sentence below was checked against the issuing body rather than written from memory, on 9 August 2026. Vendor material is labelled as vendor-reported throughout and is never presented as independent verification. Where two primary sources conflict, both are noted. Where I could not verify something, it is not asserted.

One structural caution about the subject itself. Several of the frameworks cited here changed within the last twelve months in ways that invalidate guidance still in wide circulation — among them the OWASP LLM identifiers renumbered in August 2026, a new OWASP agentic list in December 2025, SLSA v1.1 retired in November 2025, the OpenSSF baseline reissued in February 2026, FedRAMP restructured through 20x, and CMMC Phase II suspended in July 2026. Parts of this page will be wrong within a year. That is a property of the material. The fix is a re-verification cadence rather than a review date.

Forward deployed engineering: primary company sources

  1. Amazon Web Services, “AWS invests $1 billion to embed AI forward deployed engineers with customers”, 30 June 2026. aboutamazon.com/news/aws/aws-1-billion-forward-deployed-ai-engineers — vendor announcement. Accessed 9 Aug 2026.
  2. Palantir Technologies, Forward Deployed Software Engineer role descriptions, official applicant tracking system. jobs.lever.co/palantir — vendor self-description. Accessed 9 Aug 2026.
  3. OpenAI, Forward Deployed Engineer and Technical Deployment Lead role descriptions. openai.com/careers — vendor self-description. Accessed 9 Aug 2026.
  4. Anthropic, Forward Deployed Engineer, Applied AI (Federal Civilian) role description. job-boards.greenhouse.io/anthropic — vendor self-description. Accessed 9 Aug 2026.
  5. Accenture, “Accenture launches Microsoft Forward Deployed Engineering practice”, 18 March 2026. newsroom.accenture.com — vendor announcement; no financial figure disclosed.
  6. OpenAI, “OpenAI launches the Deployment Company”, 11 May 2026. openai.com/index/openai-launches-the-deployment-company — vendor announcement.
  7. Anthropic, enterprise AI services company announcement, 4 May 2026, subsequently branded July 2026. anthropic.com/news/enterprise-ai-services-company — vendor announcement. Note it uses “Applied AI engineers” rather than the forward deployed label.
  8. Databricks, AI Engineer — FDE role description, listed under Professional Services Operations. databricks.com/company/careers — vendor self-description. Accessed 9 Aug 2026.
  9. Vercel, Forward-Deployed Engineer role description. vercel.com/careers — vendor self-description. Accessed 9 Aug 2026.
  10. Salesforce, “Forward deployed engineer”, company blog, 19 November 2025. salesforce.com/blog/forward-deployed-engineer — vendor-reported; the thousand-engineer figure is a stated commitment.

Independent and analyst commentary

  1. Marty Cagan, “Forward Deployed Engineers”, Silicon Valley Product Group, 17 September 2025. svpg.com/forward-deployed-engineers — independent practitioner advocacy, not empirical research.
  2. Jina Yoon, “Forward deployed engineer”, PostHog, 11 February 2026. posthog.com/blog/forward-deployed-engineer — company blog. The on-site and infrastructure-access passage describes the industry pattern; PostHog’s own engineer is fully remote and asynchronous.
  3. Gergely Orosz, “Forward deployed engineers” (12 August 2025) and “The Pulse: forward deployed engineering heats up again” (24 May 2026), The Pragmatic Engineer. newsletter.pragmaticengineer.com — independent journalism. Contains no economic data; do not cite for unit economics.
  4. Peter Bendor-Samuel, “When are forward deployed engineers essential, and when are they not?”, Forbes / Everest Group, 30 April 2026. forbes.com — analyst-firm commentary. Contains no cost or margin data.
  5. Rocio Wu, “The uncomfortable truth about FDEs”, F-Prime Capital, 6 February 2026. fprimecapital.com/blog/the-uncomfortable-truth-about-fdes — venture-capital commentary. The 30–40% threshold is a stated heuristic, not a measured finding.
  6. Larry Dignan, “AWS launches forward deployed engineering unit”, Constellation Research, 30 June 2026. constellationr.com — analyst commentary, quoting R “Ray” Wang.
  7. Joe Schmidt, “Trading margin for moat”, Andreessen Horowitz, 4 June 2025. a16z.com/services-led-growth — venture-capital advocacy; a16z invests in companies using this model.
  8. “Forward Deployed Engineer”, Wikipedia. en.wikipedia.org/wiki/Forward_Deployed_Engineer — used only as a signpost to primary sources. Thinly sourced; several of its comparisons are editorial synthesis rather than sourced company statements.
  9. Financial Times reporting on growth in forward deployed engineering job postings, as relayed by PYMNTS (10 March 2026) and Fast Company (5 November 2025). The underlying dataset is not identified in any source located; the figure counts postings rather than filled roles.
  10. Richard MacManus, interview coverage including Sierra on the ambiguity of the term, Latent.Space, 1 July 2026. latent.space — independent technology journalism.

Standards, government and regulatory

  1. NIST SP 800-207, Zero Trust Architecture, August 2020. csrc.nist.gov/pubs/sp/800/207/final — final, not superseded.
  2. NIST SP 1800-35, Implementing a Zero Trust Architecture, NCCoE, final 10 June 2025. csrc.nist.gov/pubs/sp/1800/35/final — the earlier lettered draft volumes are withdrawn.
  3. NIST CSWP 29, The NIST Cybersecurity Framework (CSF) 2.0, 26 February 2024. csrc.nist.gov/pubs/cswp/29 — final.
  4. NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations, control catalog Release 5.2.0, 27 August 2025. csrc.nist.gov/pubs/sp/800/53/r5/upd1/final — cite the release as well as the revision.
  5. NIST SP 800-37 Rev. 2, Risk Management Framework for Information Systems and Organizations. nvlpubs.nist.gov — source of the assessment and authorization language quoted in section 22.
  6. NIST SP 800-53A Rev. 5, Assessing Security and Privacy Controls. nvlpubs.nist.gov
  7. NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0), 26 January 2023. nist.gov/itl/ai-risk-management-framework — final; NIST states a revision is in progress. No public draft had been issued as of 9 August 2026.
  8. The White House, Winning the Race: America’s AI Action Plan, July 2025. whitehouse.gov — directs the AI RMF revision; cited as the reason the framework is under revision.
  9. NIST AI 600-1, AI RMF: Generative Artificial Intelligence Profile, 26 July 2024. nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf — source of the “confabulation” terminology.
  10. NIST AI 100-2e2025, Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations, March 2025. nvlpubs.nist.gov — treats indirect prompt injection as an open problem area.
  11. ISO/IEC 27001:2022 including Amendment 1:2024 (climate action changes). iso.org/standard/27001 — transition from the 2013 edition closed in October 2025.
  12. ISO/IEC 42001:2023, Artificial intelligence — Management system; and ISO/IEC 42006:2025, requirements for certification bodies. iso.org/standard/42001 · iso.org/standard/42006
  13. AICPA, 2017 Trust Services Criteria for Security, Availability, Processing Integrity, Confidentiality, and Privacy (with revised points of focus — 2022). aicpa-cima.com — SOC 2 is an attestation examination, not a certification.
  14. Center for Internet Security, CIS Critical Security Controls v8.1, 25 June 2024. cisecurity.org/controls/v8-1
  15. SLSA v1.2 specification, approved 24 November 2025. slsa.dev/spec/v1.2 — v1.1 is explicitly marked retired.
  16. OpenSSF Open Source Project Security Baseline, release 2026-02-19. baseline.openssf.org
  17. HHS, HIPAA Security Rule and the proposed rule to strengthen cybersecurity of ePHI (RIN 0945-AA22), published 6 January 2025. hhs.gov · federalregister.gov · reginfo.gov — not finalized; now listed under long-term actions. The existing Security Rule at 45 CFR 164 Subpart C governs.
  18. PCI Security Standards Council, PCI DSS v4.0.1, June 2024. pcisecuritystandards.org — the previously future-dated v4.x requirements have been in force since 31 March 2025.
  19. FedRAMP 20x programme and Consolidated Rules for 2026. fedramp.gov/20x — Phase 3 active from April 2026; terminology and artifacts renamed.
  20. US Department of Defense CIO, CMMC programme status including the suspension of Phase II requirements announced 13 July 2026; DFARS Case 2019-D041 final rule effective 10 November 2025. dodcio.defense.gov/CMMC · federalregister.gov
  21. NIST SP 800-171 Rev. 3, Protecting Controlled Unclassified Information in Nonfederal Systems and Organizations, 14 May 2024. csrc.nist.gov/pubs/sp/800/171/r3/final — note that CMMC Level 2 is still assessed against Rev. 2.

AI and application security guidance

  1. OWASP, LLM01 Prompt Injection, OWASP Top 10 for LLM Applications. owasp.org — source of the statement that fool-proof prevention is unclear.
  2. OWASP GenAI Security Project, OWASP Top 10 for LLM Applications 2026, August 2026. genai.owasp.org — identifiers renumbered from the 2025 edition; Excessive Agency is now third. OWASP’s own legacy page still displays 2025 identifiers.
  3. OWASP GenAI Security Project, OWASP Top 10 for Agentic Applications, December 2025 (ASI01–ASI10). genai.owasp.org
  4. OWASP MCP Top 10. owasp.org/www-project-mcp-top-10 — Incubator project at v0.1 beta; cite with less weight than a flagship document.
  5. MITRE ATLAS, data release v2026.05, May 2026. atlas.mitre.org · github.com/mitre-atlas — continuously updated; techniques now tagged by platform including agentic AI.
  6. Dane Stuckey, Chief Information Security Officer, OpenAI, public statement on agent-mode security, 21 October 2025. Vendor statement — “prompt injection remains a frontier, unsolved security problem”.
  7. Anthropic, “Prompt injection defenses”, 24 November 2025. anthropic.com/research/prompt-injection-defenses — vendor-reported; states a 1% attack success rate “still represents meaningful risk”.
  8. Greshake, Abdelnabi, Mishra, Endres, Holz & Fritz, “Not what you’ve signed up for: compromising real-world LLM-integrated applications with indirect prompt injection”, ACM AISec 2023, arXiv:2302.12173. The canonical citation for the indirect prompt injection class.
  9. Model Context Protocol specification, revision 2026-07-28, including Security Best Practices. modelcontextprotocol.io/specification — source of the quoted normative requirements.
  10. CVE-2025-49596, MCP Inspector unauthenticated remote code execution, CVSS v4.0 base score 9.4, published 13 June 2025. nvd.nist.gov/vuln/detail/CVE-2025-49596

Research

  1. Palantir, Ontology documentation and core concepts. palantir.com/docs/foundry/ontology — vendor documentation; no independent technical evaluation located.
  2. Palantir, “Reducing hallucinations with the Ontology in Palantir AIP”, 8 July 2024. Vendor engineering blog. Note the verb: reducing, paired with human oversight.
  3. Edge, Trinh, Cheng, Bradley, Chao, Mody, Truitt, Metropolitansky, Ness & Larson, “From local to global: a Graph RAG approach to query-focused summarization”, arXiv:2404.16130. Microsoft Research. Headline result concerns comprehensiveness and diversity, evaluated by preference judgement, not factual accuracy.
  4. Microsoft, GraphRAG responsible AI transparency documentation. github.com/microsoft/graphrag — states that expert human verification of answers is needed and that the system is designed for trusted users.
  5. Chauhan, Raj, Mujumdar, Saha & Jain, “Mind the query”, EMNLP 2025 Industry Track. aclanthology.org — peer-reviewed; 76.75% execution accuracy overall for the strongest model tested, 49.93% on complex aggregation.
  6. Ranganath & Raghavendra, “PIPE-Cypher”, arXiv:2606.08481, June 2026. Preprint, not peer-reviewed. Reports 0.916 schema validity against 0.189 exact execution accuracy.
  7. Shen, Wan et al., “Understanding, detecting and repairing real-world in-context-learning-based text-to-SQL errors”, PACM SE (FSE), arXiv:2501.09310. Peer-reviewed; semantic errors account for 36.1% of errors on one benchmark.
  8. BIRD text-to-SQL benchmark leaderboard. bird-bench.github.io — accessed 9 Aug 2026; human baseline 92.96% execution accuracy.
  9. Niu et al., “RAGTruth: a hallucination corpus for developing trustworthy retrieval-augmented language models”, ACL 2024. aclanthology.org — peer-reviewed.
  10. Jacovi et al., “The FACTS Grounding leaderboard”, arXiv:2501.03200. Google DeepMind — vendor-reported research.
  11. Joren, Zhang, Ferng, Juan, Taly & Rashtchian, “Sufficient context: a new lens on retrieval augmented generation systems”, ICLR 2025, arXiv:2411.06037. Peer-reviewed.
  12. Xu, Jain & Kankanhalli, “Hallucination is inevitable: an innate limitation of large language models”, arXiv:2401.11817. Preprint, widely cited and contested.
  13. Kalai, Nachum, Vempala & Zhang, “Why language models hallucinate”, arXiv:2509.04664; OpenAI, 5 September 2025. Vendor-reported. Argues confident hallucination is reducible through abstention while accuracy never reaches 100%.
  14. Huang, Chen, Mishra, Zheng, Yu, Song & Zhou, “Large language models cannot self-correct reasoning yet”, ICLR 2024, arXiv:2310.01798. Google DeepMind; peer-reviewed.
  15. Peeters, Bizer et al., “WDC Products: a multi-dimensional entity matching benchmark”, EDBT 2024, arXiv:2301.09521. Peer-reviewed; 89.04 F1 on seen entities against 64.56 on unseen.
  16. Agrawal, Kumarage, Alghamdi & Liu, “Can knowledge graphs reduce hallucinations in LLMs? A survey”, NAACL 2024, arXiv:2311.07914. Peer-reviewed; frames the contribution as mitigation.
  17. Ozsoy, “Enhancing Text2Cypher with schema filtering”, arXiv:2505.05118. Neo4j — vendor-reported research; benefit was not uniform across models.
  18. METR, “Measuring the impact of early-2025 AI on experienced open-source developer productivity”, 10 July 2025, arXiv:2507.09089. Randomised controlled trial, 16 developers, 246 issues. A snapshot of early-2025 tooling on a specific kind of work.
  19. Zhu, Jin, Pruksachatkun et al., “Establishing best practices for building rigorous agentic benchmarks”, arXiv:2507.02825. Preprint.
  20. Liang, Bommasani, Lee et al., “Holistic evaluation of language models (HELM)”, TMLR 2023, arXiv:2211.09110. Peer-reviewed; seven evaluation axes.
  21. Ong, Almahairi, Wu, Chiang, Wu, Gonzalez, Kadous & Stoica, “RouteLLM: learning to route LLMs with preference data”, ICLR 2025, arXiv:2406.18665. Peer-reviewed; reported savings vary by an order of magnitude across benchmarks.
  22. Yan, Bischof, Frye, Husain, Liu & Shankar, “What we learned from a year of building with LLMs”, O’Reilly Radar, 28 May 2024. Practitioner synthesis, vendor-neutral.
  23. He, “Defeating nondeterminism in LLM inference”, Thinking Machines Lab, 10 September 2025. Vendor-reported but reproducible; 1,000 completions at temperature zero produced 80 unique outputs before batch-invariant kernels were applied.
  24. Turan, “Oversight has a capacity: calibrating agent guards to a subjective, fatiguing human”, arXiv:2606.08919, June 2026. Single-author preprint, not peer-reviewed. Cited for direction, not for its specific figures.
  25. Parasuraman & Riley, “Humans and automation: use, misuse, disuse, abuse”, Human Factors 39(2), 1997. The canonical source for automation misuse and complacency; cited for the concept.

Discuss enterprise architecture. If your organization is standing up a forward deployed capability, or working out whether an AI deployment is ready to be called production, I am happy to talk through how this reference translates to your environment. Get in touch or explore more projects.