Forward Deployed Engineering
A reference architecture and operating model for taking complex enterprise and AI systems from prototype to governed production — and an account of where the model breaks when nobody is watching for it.
What this is. I wrote this reference architecture to work out how forward deployed engineering actually moves complex AI and enterprise systems from prototype into governed production, and where the operating model and the architecture supporting it tend to fail. It is authored research. It does not describe a client engagement, and no deployment, customer, saving or audit result is claimed anywhere on this page. The four scenarios in Reference scenarios are composites assembled from common industry constraints and are labelled as such.
The source material I started from contained several claims that did not survive verification, including a specific financial saving, a regulatory audit outcome and the assertion that a knowledge graph produces zero hallucinations. I have published those corrections rather than quietly dropping them. See Corrections.
01 The deployment gap
A vendor sells a capability. A customer needs an outcome. Between those two things sit five boundaries, and most implementations fail at one of them rather than at the technology itself.
Nothing about that is new. What has changed is the volume and the stakes: enterprises are now buying systems whose behaviour is probabilistic, whose integration surface reaches into their most sensitive data, and whose value depends entirely on people changing how they work. AWS announced a billion-dollar commitment to a dedicated forward deployed engineering organization in June 2026 on exactly this reasoning.[1] Accenture announced a Microsoft-focused forward deployed engineering practice three months earlier, on its own account.[5] OpenAI and Anthropic each spun delivery into a separate services venture rather than scaling it inside the lab.[6][7]
Select a boundary to see what actually goes wrong there.
capability
systems
workflows
outcomes
Each of these is solvable. None of them is solvable by documentation, and only the first is reliably solvable by a pre-sales architect who leaves after the design review. That is the argument for forward deployment, and it is worth stating in its weakest defensible form rather than its strongest: some problems sit in the seams between systems and teams, and the seams are where nobody has authority to change code.
02 What makes engineering “forward deployed”
Forward deployed engineering is an operating model, not an architecture. The architecture on this page is a synthesized reference design for supporting that model. There is no canonical forward deployed architecture, and I have found no standards body that defines one.
Five properties separate it from adjacent ways of working. An engagement missing several of them is doing something else, whatever it is called, and the label then obscures rather than clarifies what is being bought. Marty Cagan’s argument for the model is that it accelerates product discovery, and that it “applies much more broadly than just in the most challenging product situations” — worth weighing as advocacy from a product-management perspective rather than as evidence.[11]
01 Customer embedding
The engineer works close enough to users and systems to observe operational reality rather than requirements. Embedding is operational, not geographic. Palantir’s own posting for the role it originated caps travel at 25%,[2] and PostHog runs its forward deployed engineer fully remote and asynchronous.[12]
02 Engineering authority
The role can architect and write production software, not only recommend that someone else does. OpenAI describes its forward deployed engineers as owning “discovery, technical scoping, system design, build, and production rollout”.[3] Without the authority to build, the role becomes advisory work under a more energetic title.
03 End-to-end ownership
Discovery, scoping, architecture, implementation, evaluation, rollout and stabilization stay connected to the same people. Every handoff between those stages loses the context that made the previous stage correct.
04 Field-to-product feedback
Deployment problems that recur become inputs to platform and product work. This is the property most often claimed and least often instrumented, and its absence is what turns the model into custom consulting with a software company’s cost base.
05 Outcome orientation
Success is measured by production adoption and workflow effect, not by configuration completed or milestones closed. This is also where the model is most easily faked, because adoption metrics are easy to define generously.
— What it is not
It is not a security discipline, a compliance mechanism or an architecture style. It is a way of organizing delivery. Everything security-relevant on this page comes from controls that would be required regardless of who was writing the code.
The label is contested, and that is worth knowing
The term expanded very quickly. Job postings for forward deployed roles were reported by the Financial Times to have risen more than 800% between January and September 2025, a figure repeated widely although the underlying dataset is not identified in any source I could find.[19] That growth has pulled in work that does not match the definition. Constellation Research’s Ray Wang put it bluntly: “Some of these FDEs don’t necessarily fit the definition and are glorified sales and customer success people.”[16] Gergely Orosz, reading a major cloud vendor’s posting for the role, translated it as “you are a contractor who codes at a customer’s office”.[13] Sierra’s own account concedes that the term “can mean so many things” that it risks becoming meaningless.[20] Even the encyclopaedia entry for the role is thin, resting largely on vendor self-description.[18]
The vendors are not consistent either. AWS positions its practice explicitly against traditional consulting.[1] Databricks files its AI forward deployed engineering team under Professional Services Operations and describes the work as “professional services engagements”.[8] Vercel’s posting says “this is not a traditional consulting role” while reporting the role to a Director of Professional Services.[9] None of that makes the model illegitimate. It does mean that “we use forward deployed engineers” carries almost no information on its own, and that an enterprise buying the model should ask what specifically it is buying.
03 FDE and adjacent roles
The boundaries between these roles vary by company. What follows is a comparison of common operating patterns, not a set of formal industry definitions, and any given organization will place at least one of them differently.
Show all seven roles side by side
Scrollable on narrow screens. The single-role view above carries the same content in a form that reads on a phone.
| Dimension | FDE | Solutions Architect | Solutions Engineer | Professional Services | Customer Engineer | Product Engineer | TAM |
|---|---|---|---|---|---|---|---|
| Primary objective | Production outcome in one environment | A defensible design | Technical validation of a deal | Bounded delivery of a known pattern | Technical adoption and enablement | Capability for many customers | Sustained account health |
| Lifecycle stage | Post-sale through operations | Pre-sale and early design | Pre-sale | Post-sale, fixed scope | Pre- and post-sale | Continuous roadmap | Ongoing |
| Customer proximity | Embedded in the workflow | Periodic, workshop-based | Meeting-based | Project-based | Regular, advisory | Indirect | Relationship-based |
| Coding depth | Production code, substantial | Reference material, prototypes | Demos and proofs of concept | Configuration and integration code | Samples and accelerators | Product code | Little to none |
| Production accountability | Yes, through stabilization | No | No | To acceptance criteria | Shared, informal | For the platform, not the deployment | Escalation ownership |
| Customization | Deep, environment-specific | Design-level only | Illustrative | Within a defined catalogue | Light | None — generalizes instead | None |
| Product feedback duty | Explicit and instrumented | Informal | Competitive and feature gaps | Delivery friction | Adoption blockers | Receives it | Account-level themes |
| Operational ownership | Until handoff is proven | None | None | Until acceptance | None | Platform SLOs | Coordinates, does not operate |
| Typical duration | Months, occasionally longer | Weeks | Days to weeks | Fixed-term project | Ongoing, part-time | Continuous | Contract lifetime |
On the titles themselves. “Forward Deployed Software Engineer” and “Forward Deployed Security Engineer” are both attested at named employers.[2][3] “Technical Deployment Lead” appears at OpenAI, always compounded with the forward deployed engineering function rather than standing alone.[3] I could not attest “Forward Deployed Platform Engineer” as a title in use at any employer, so I am not presenting it as one; the nearest real examples are Sierra’s Forward Deployed Infrastructure Engineer and Cognition’s Deployed Engineer, which is not the same phrase. “Customer Engineer” predates this wave as a pre-sales title at Google Cloud and is not a synonym.
04 Forward-deployed reference architecture
Six runtime layers with two concerns that cut across all of them. This is a reference design I assembled to reason about the model, not a description of any vendor’s product. Named technologies are examples of what can occupy a slot, never requirements.
Select a layer for its purpose, the split of responsibility between vendor and customer, what tends to go wrong there, and example technologies.
What this diagram deliberately does not require. Kubernetes, a knowledge graph, an ontology, microservices, edge computing, serverless, GraphRAG, a data mesh, a particular cloud and a particular model vendor are all absent from the layer definitions. Each is an implementation choice that some deployments justify and most do not. Start from the capability and the constraint; introduce the technology when you can say what it is for.
05 One operating model, different deployment boundaries
The operating model barely changes across these five. What changes is where the trust boundary falls, and every consequential decision on the page follows from that: how software is updated, how the engineer reaches the system, whether telemetry ever arrives, and who holds the credentials.
No platform has to support all five. Supporting the air-gapped case imposes design constraints on everything else — offline licensing, signed offline bundles, deferred telemetry, local policy decision points, an update path that survives months of disconnection. Building those for a customer base that is entirely cloud-resident is a large, permanent tax on velocity. Deciding which topologies you will not support is an architecture decision and should be recorded as one.
06 Decisions that matter
These are the ten decisions I have found determine whether a forward deployed engagement stays maintainable. Each has a genuine case on both sides; the failure is not choosing wrong, it is choosing implicitly and discovering the choice two years later during an upgrade.
Centralized data versus federated query Data
Drivers. Data residency and sovereignty rules, cross-border transfer mechanisms, the sensitivity of the underlying records, query latency tolerance, and whether the analytical questions are known in advance.
Centralize when the analytical workload is exploratory, the data can lawfully move, and join performance across the whole corpus matters more than locality.
Federate when raw records cannot cross a jurisdictional or organizational boundary. Push the query to the data, apply local masking and aggregation, and return only results. The cost is real: federated queries are slower, harder to optimize, harder to debug, and each local node becomes a component you have to operate and patch.
Risk if implicit. Teams build a central lake, then discover a residency constraint, then bolt on regional exceptions until the model is neither central nor federated and nobody can describe where a given record lives.
Cloud inference versus local inference AI
Drivers. Data sensitivity, latency budget, connectivity, unit cost at expected volume, model capability required, and how often the model needs replacing.
Cloud when you need frontier capability, volumes are variable, and the data can leave the estate under an acceptable contractual and technical control set.
Local when connectivity is constrained or absent, latency is measured in tens of milliseconds, or the data cannot leave. Accept the consequences: a smaller model, a deployment pipeline for model artifacts, drift monitoring you have to build, and a hardware refresh problem.
The middle path that usually wins is decoupling training from inference — train centrally, ship a compact evaluated artifact to the edge, run drift detection locally, and queue retraining requests for whenever the link returns.
Product configuration versus custom extension Delivery
Drivers. How far the requirement is from the product’s intended use, how many other customers would want it, who will maintain it in three years, and whether the platform has an extension contract worth depending on.
Configure when the gap is a setting, a policy or a template. Configuration survives upgrades; code does not, unless someone maintains it.
Extend when the requirement is genuinely specific and the extension point is versioned and tested. Register the extension, name its owner, record which platform API version it depends on, and set a disposition date.
Risk if implicit. Custom code written under delivery pressure with no owner and no registry entry is, in my reading, the most consequential source of upgrade paralysis in this model.
Synchronous versus asynchronous integration Integration
Drivers. Whether the caller can wait, the availability of the downstream system, whether the operation is idempotent, and how failures should surface to a human.
Synchronous when the user is waiting for a decision and the downstream system has a credible availability target. Bound the timeout, define the fallback, and decide in advance what the interface says when the dependency is down.
Asynchronous when the downstream is a legacy system with batch semantics, the operation is expensive, or partial availability is normal. The cost is that you now own a queue, a retry policy, a dead-letter path, an idempotency key strategy and a reconciliation process.
Retrieval over documents versus retrieval over a graph AI
Drivers. Whether the questions are relationship-heavy, whether the entities are resolvable, whether provenance must be traceable to a specific record, and whether anyone will maintain the schema.
Document retrieval when the corpus is textual, the questions are local, and the answer lives in a passage. It is cheaper, simpler and has a shorter failure chain.
Graph retrieval when answering requires traversing relationships between resolved entities, when authorization needs to be expressed over those relationships, or when provenance must point at identified nodes and edges rather than at a chunk of text.
What neither buys you. Correctness. See GraphRAG and its limits for the measured numbers, which are worse than most architecture diagrams imply.
Direct model access versus an AI gateway AI
Drivers. Number of applications, need for central policy, cost attribution, the likelihood of changing model vendor, and whether anyone needs a single audit trail of model use.
Direct when there is one application, one model and one team. A gateway for a single caller is infrastructure you maintain for no benefit.
Gateway when you need central authentication, per-tenant quota and cost attribution, prompt and model version pinning, uniform logging, and the ability to substitute a model without touching applications. The gateway becomes a dependency on the critical path and needs its own availability target and failure mode.
Autonomous agent action versus human-approved action AI
Drivers. Reversibility of the action, blast radius, the cost of a false positive against the cost of delay, and how many approvals a reviewer will see per day.
Autonomous when the action is reversible, bounded and cheap to get wrong — drafting, classifying, enriching, proposing.
Human-approved when the action moves money, changes access, deletes data, contacts a customer or cannot be undone. Approval is a real control that degrades under volume, so it has to be rationed. Requiring approval for everything produces a queue that gets cleared, not read.[75] Automation-induced complacency is a long-established human-factors finding rather than a new observation about AI.[76]
The control that does not depend on attention is authorization enforced in the downstream system. OWASP’s own guidance is explicit: “implement authorization in downstream systems rather than relying on an LLM to decide if an action is allowed”.[43]
Shared vendor plane versus customer-isolated runtime Platform
Drivers. Regulatory posture, tolerance for shared-tenancy risk, the customer’s own security review process, and what the vendor can afford to operate.
Shared when the customer accepts multi-tenant isolation and the vendor can demonstrate it. Operating cost per customer stays low and upgrades are uniform.
Isolated when a regulator, a contract or a threat model requires it. Every isolated runtime is a separate environment to patch, monitor, back up and prove. Ten isolated runtimes multiply the operational surface tenfold; revenue rarely follows the same curve.
Standing remote support access versus customer-controlled just-in-time access Security
Drivers. Incident response time, the customer’s privileged access model, audit requirements, and whether the vendor is treated as a third party or a sub-processor.
Standing access shortens mean time to repair and is what most engagements drift into. It also means a vendor credential compromise is a customer production incident, and it is the finding an assessor will write up first.
Just-in-time access costs minutes at the start of every incident and removes an entire category of standing risk. Requested through the customer’s own workflow, scoped to a role, time-boxed, session-recorded, reviewed. Design for this from the first day, because retrofitting it after two years of standing access is an organizational fight, not a technical one.
Customer-specific code versus product primitive Product
Drivers. How many customers need it, whether it exposes a missing platform capability, who will own its lifecycle, and whether it can be tested and versioned independently.
Keep it specific when the requirement encodes one customer’s process. Promoting local process into a shared platform makes every future customer carry it.
Promote to a primitive when the same shape has appeared three times, the abstraction is stable, and moving it into the core reduces total delivery cost across the portfolio rather than just in the next engagement.
The intermediate step people skip is an extension contract — a supported, versioned point where the specific thing can live without becoming core. See Productization economics.
07 Engagement lifecycle
Eleven stages. The first one is the one most organizations skip, and skipping it is why the model gets a reputation for being expensive.
Deliberately without a calendar. The source material I worked from offered two adoption timelines for structurally similar programmes that differed by a factor of two to three, neither attributed. Published month ranges for work like this are estimates that a sponsor will hold you to. What is durable is the ordering and the exit criteria: qualification before discovery, evaluation before productionization, readiness proven before rollout, handoff proven before the engagement is called finished.
08 The feedback loop
This is the part of the model that determines whether it is a business or a service line with unusually expensive staff.
The model scales only when field learning improves the platform. If every engagement stays bespoke, forward deployment is custom consulting carrying a software company’s valuation. F-Prime Capital, writing from an investor’s perspective, proposes a working threshold: if more than 30–40% of deployments require significant forward deployed effort, the problem has stopped being go-to-market and become product design.[15] That figure is a practitioner heuristic rather than a measured finding, and I would treat it as a prompt to instrument rather than a number to manage to. The underlying claim is harder to argue with: “if every deal requires bespoke engineering, you don’t have a product; you have a consultancy with a logo”.[15]
There is a serious counter-position. Andreessen Horowitz argues that optimizing for gross margin percentage is the wrong objective early, and points at ServiceNow and Workday, both of which entered public markets in the fifties and low sixties and reached the mid-to-high seventies years later.[17] Both parties are investors and both have a position to talk. The reconciliation I find defensible is narrow: services-heavy delivery is a legitimate investment when it is buying a repeatable capability, and it is a permanent cost when it is buying a customer. The instrumentation that tells you which one is happening is in Measuring outcomes, and most organizations running this model do not have it.
09 Productization economics
Every forward deployed engagement produces code that should not stay where it was written. The discipline is knowing which code that is, and having somewhere for it to go that is not the core product.
The promotion questions
Before customer-specific work is promoted toward the core, these nine questions decide whether it should be. A “no” to the last one is disqualifying regardless of the others.
- How many customers actually need this, as opposed to how many might?
- Is the requirement domain-specific, or does it encode one organization’s internal process?
- Does it expose a platform primitive that is missing, or does it work around one that exists?
- What permanent complexity does supporting it add, and to whom?
- Who owns its lifecycle after the engagement ends, by name?
- Can it be tested independently of the deployment it came from?
- Can it be versioned, deprecated and removed?
- Can it live outside the core behind an extension contract instead?
- Does moving it into the core reduce total delivery cost across the portfolio, rather than only in the next engagement?
The customization debt register
Every customization that is not immediately promoted or retired goes into a register. This is the artifact that makes upgrade planning possible, and its absence is why organizations discover during a major version upgrade that they cannot enumerate what they have built.
| Field | Why it is there |
|---|---|
| Customization | What was built and where the code lives. A repository link, not a description. |
| Customer dependency | Which deployments would break if it were removed. Determines whether removal is a decision or a negotiation. |
| Owner | A named engineer and a named team. A team alias is not an owner. |
| Reusable potential | Assessed against the nine promotion questions, not by enthusiasm. |
| Core API dependency | Which platform interfaces and which versions. This is the field that makes blast-radius analysis possible before an upgrade. |
| Upgrade risk | What breaks at the next major version, assessed rather than assumed. |
| Maintenance effort | Observed, not estimated. Effort that nobody measures gets attributed to delivery and disappears. |
| Target disposition | Promote, keep as an extension, replace with product configuration, or retire — with a date. An entry with no disposition stays forever, because nothing ever forces the decision. |
10 When the product is AI
Forward deployment of a deterministic application is a hard integration problem. Forward deployment of an AI system is that plus a component whose behaviour is part of the production system and changes when the vendor ships a model. Everything below exists because of that difference.
Two claims are worth stating before the diagram, because architectures in this space are frequently built on their opposites. Model capability is not system reliability: a model that answers well in a demo is evidence about the demo. And benchmark performance does not predict workflow performance — a randomised controlled trial published by METR in 2025 found experienced open-source developers took 19% longer to complete real issues when allowed to use AI tooling, while predicting a 24% speedup beforehand and still believing they had been 20% faster afterwards.[69] That result is a snapshot of early-2025 tooling on a specific kind of work and should not be generalized further than that. What generalizes is the shape of the error: the people closest to the system were confidently wrong about its effect, in the same direction, and only measurement caught it.
Select a stage of the production path for the engineering concerns that attach to it.
Prompt injection is a design constraint, not a pending fix
An AI system that reads untrusted content and can take actions has an attacker-controlled input path into its instructions. This is not a solved problem and no credible source claims it is. OWASP’s LLM Top 10 states it directly: “given the stochastic influence at the heart of the way models work, it is unclear if there are fool-proof methods of prevention for prompt injection”.[42] OpenAI’s Chief Information Security Officer described it in October 2025 as “a frontier, unsolved security problem”.[47] Anthropic, publishing a defence that reduced attack success to 1% against an adaptive attacker, wrote that this “still represents meaningful risk” and that they were sharing the results “to demonstrate progress, not to claim the problem is solved”.[48] The attack class was formally characterized in 2023 and has not been closed since.[49] NIST’s adversarial machine learning taxonomy treats it as an open problem area and warns that “many of these mitigations may themselves be vulnerable to new discoveries and evolutions in attacker techniques”.[30]
The architectural consequence is specific. Split your controls into two sets and be honest about which is which. Input filtering, provenance framing and content classification reduce the rate at which injection succeeds; they are classifiers operating on adversarial natural language and cannot carry a safety case. Authorization enforced in the downstream system, identity and audience binding, egress control and approval bound to a specific payload do not care whether the model was fooled. Put the safety case on the second set and budget the first as defence in depth.
Tool integration and MCP
Where a system reaches tools through the Model Context Protocol or an equivalent, the protocol specification itself is unusually clear about what it does not do. Revision 2026-07-28 states that “descriptions of tool behavior such as annotations should be considered untrusted, unless obtained from a trusted server”, and that “while MCP itself cannot enforce these security principles at the protocol level, implementors SHOULD” build the controls themselves.[50] It also carries hard normative requirements that early integrations routinely violated, including “MCP servers MUST NOT accept any tokens that were not explicitly issued for the MCP server” and “MCP servers MUST NOT treat possession of a state handle as authentication”.[50]
The risk is not theoretical. CVE-2025-49596 in the MCP Inspector developer tool carried a CVSS v4.0 base score of 9.4 for unauthenticated remote code execution caused by a missing authentication step between the inspector client and its proxy.[51] A forward deployed engagement that connects an agent to customer systems inherits every one of these decisions. The minimum set worth insisting on: an inventory of which servers are trusted and who approved them, per-target token exchange with an explicit audience rather than passthrough, tool permissions scoped to the single call, secrets held outside the tool’s reach, validation of the action rather than of the intent behind it, transaction and rate limits, and an audit record that survives the compromise of the thing being audited.
OWASP published a Top 10 for Agentic Applications in December 2025 with its own identifier series — ASI01 Agent Goal Hijack through ASI10 Rogue Agents — which is a better fit for this material than the LLM list alone.[44] There is also an OWASP MCP Top 10, but it is an Incubator project at version 0.1 and should not be cited with the weight of a flagship document.[45]
11 GraphRAG: useful grounding, not a guarantee
A semantic layer over enterprise data is a good idea for reasons that have nothing to do with hallucination. It is also routinely sold on a claim that the measurements do not support, and the correction matters because the claim changes what controls people think they need.
The claim I am correcting. One of the source documents behind this project states that “to achieve zero hallucinations, the LLM is forced to write a formal graph database query… based only on the schema”, and that integrating an LLM with an ontology “shifts AI architecture from probabilistic guessing to deterministic reasoning”. Neither is true, and the second document in the same family contradicts the first by describing a verification agent whose existence presupposes that the first stage produces errors.
Notably, the vendor most associated with this architecture does not make the claim either. Palantir’s own engineering post on the subject is titled Reducing Hallucinations with the Ontology, describes the ontology as helping to ground responses, and pairs it with human oversight.[53] The peer-reviewed literature uses the same verb: a NAACL 2024 survey of knowledge graphs and LLMs frames the contribution as mitigating hallucination, not removing it.[67]
What a graph or ontology genuinely provides
- Retrieval constrained to entities and relationships that exist, which removes a class of fabricated references.
- Provenance that points at identified nodes and edges rather than at a text chunk, which is the strongest governance argument for the pattern.
- A structured place to enforce authorization, including permissions inherited through relationships rather than restated per object.
- Better handling of relationship-heavy questions than passage retrieval, which is what graph structures are actually for.
- Normalization across heterogeneous sources, with the trade-off stated: normalizing an EC2 instance and an Azure VM into one concept discards the provider-specific attributes that some questions need.
- An execution step that is genuinely deterministic. Given a fixed query and a fixed graph state, the result is reproducible and auditable.
What it does not provide
The pipeline has three stages and only the middle one is deterministic. Translating a question into a query is a language-model operation. Turning the returned rows back into prose is a language-model operation. Both are probabilistic, and constraining the output to a valid schema does not change that — it changes the failure mode from fabricated prose to a query that is syntactically valid, schema-conformant, and answers a different question than the one asked, returning a clean-looking result set with no signal that anything went wrong.
The numbers
These are the measurements that made me rewrite this section rather than soften it.
| Measurement | Result | Source |
|---|---|---|
| Text-to-Cypher execution accuracy | 76.8% overall for the strongest model tested; 49.9% on complex aggregation | Purpose-built Cypher benchmark, EMNLP 2025 Industry Track[56] |
| Schema validity against execution accuracy | 0.916 schema-valid, 0.189 correct answers — same model, same tasks | Enterprise-schema evaluation, 2026 preprint[57] |
| Text-to-SQL, the mature analogue | Best systems around 82% execution accuracy against a human baseline of 92.96% | BIRD benchmark leaderboard[59] |
| Errors that execute cleanly | 36.1% of errors on one benchmark are semantic faults that run without error and return wrong data | Error taxonomy, PACM SE / FSE[58] |
| Entity resolution on unseen entities | 89.0 F1 on entities seen in training, 64.6 F1 on unseen — roughly a 25-point drop | WDC Products benchmark, EDBT 2024[66] |
| Hallucination with retrieval present | 43.1% of responses in one RAG corpus contained at least one hallucination | RAGTruth, ACL 2024[60] |
| Groundedness when the source is in the prompt | Best systems in the low-to-mid 80s%; nothing reaches 100% | FACTS Grounding, Google DeepMind — vendor-reported[61] |
The 0.916 against 0.189 pair is the one to keep. A schema checker passing more than nine in ten generated queries while fewer than one in five return the right answer is the arithmetic of why schema verification is not error elimination. It catches references to labels and properties that do not exist, which are the cheapest errors. The expensive ones use the right vocabulary to ask the wrong question.
Entity resolution deserves separate attention because a knowledge graph inherits every merge error and every missed merge as a fact. The benchmark drop from 89.0 to 64.6 F1 falls precisely on the case an enterprise graph meets every time a new supplier, customer or asset appears.[66] That is not an argument against building the graph. It is an argument for treating it as a maintained asset with a measured resolution quality rather than as ground truth.
Nor does an LLM verification step close the gap on its own. Google DeepMind’s ICLR 2024 result is that language models “struggle to self-correct their responses without external feedback, and at times, their performance even degrades after self-correction”.[65] A schema is external feedback about schema errors, which is why that check works. It is not external feedback about whether the query captured the user’s intent.
Mitigations that actually apply
01 Schema-constrained generation
Present a pruned, relevant schema rather than the whole model. Neo4j’s own research found that reducing schema size improved accuracy for most models tested and reduced cost for all of them, though not uniformly and from a low base.[68]
02 Query validation before execution
Reject references to labels, properties and relationships that do not exist. Cheap, worth having, and explicitly insufficient on its own.
03 Access-aware retrieval
Evaluate the requesting principal’s permissions at the enforcement point, not by trusting a claim the model passed along. Retrieval that ignores authorization turns a search feature into a data exfiltration path.
04 Deterministic rules where they exist
If a business rule can be expressed in code, express it in code. Asking a model to apply a rule you could have written is choosing a probabilistic implementation of a deterministic requirement.
05 Result validation
Check the returned set against expectations that do not come from the model — cardinality, type, range, referential consistency. This is the only automatic check positioned to catch the right-vocabulary-wrong-question failure.
06 Abstention as a first-class outcome
OpenAI’s position is that hallucination persists partly because “standard training and evaluation procedures reward guessing over acknowledging uncertainty”.[64] Build a path for “I cannot answer this from the available data” and evaluate it as a success.
07 Human review for consequential actions
Rationed to actions that are irreversible or high-value, because approval degrades under volume and an approval queue that is always full stops being read.
08 Continuous evaluation
Retrieval quality and generation quality measured separately. ICLR 2025 work found models fail differently depending on whether the retrieved context was sufficient, and that insufficient context is worse than none — one model’s incorrect-answer rate rose from 10.2% with no context to 66.1% with insufficient context.[62]
On GraphRAG specifically. Microsoft’s GraphRAG is worth reading for what its authors claim, which is narrower than how it is usually cited. The paper’s headline result is improvement in comprehensiveness and diversity of answers to global sensemaking questions over a corpus, evaluated by preference judgement, not factual accuracy.[54] The project’s own responsible-AI documentation states that “human analysis by a domain expert of the answers is needed in order to verify and augment GraphRAG’s generated responses” and that “the system is designed for trusted users”.[55] Hallucination there is measured as a rate.
12 Zero Trust is broader than data access
Access control at the semantic layer is a real Zero Trust control and a good one. It is not Zero Trust, and the difference is not pedantry — it decides which controls a design review will find missing.
NIST SP 800-207 defines zero trust architecture as “an enterprise’s cybersecurity plan that utilizes zero trust concepts and encompasses component relationships, workflow planning, and access policies”, and states plainly that “ZT is not a single architecture but a set of guiding principles” and that transitioning to it “cannot simply be accomplished with a wholesale replacement of technology”.[21] The document lists seven tenets. Role- and attribute-based authorization enforced at a data or ontology layer contributes strongly to three of them — per-session access decisions, dynamic policy incorporating identity and asset state, and authorization strictly enforced before access is allowed. It contributes nothing on its own to the other four: treating all data sources and computing services as resources, securing all communication regardless of network location, continuously measuring the integrity and security posture of every asset, and collecting state telemetry to improve posture.
The practical consequence for a forward deployed engagement is that the security conversation cannot be delegated to whoever owns the data model. Workload identity, secrets custody, supply chain provenance and the telemetry that makes continuous verification possible all sit with different teams, usually on the customer side, and all of them need to be in the architecture before the first production deployment rather than after the first assessment.
13 The access boundary
This is the diagram I would put in front of a customer security team first, because it answers the question they are actually asking: what can your engineer do in my production environment, and how would I know?
Two things make this diagram hold up in practice rather than only in a document. Audit custody sits outside the administrative domain of the people being audited, so the evidence survives the case where the vendor account is the thing that was compromised. And access is requested through the customer’s own workflow rather than the vendor’s, which is the difference between the customer being able to revoke access and the customer being told that access was revoked.
Everest Group makes the sharper version of the point, and it is worth taking seriously precisely because it comes from outside the enthusiasm: introducing forward deployed engineers into an environment with mature change control “can bypass these safeguards”, and “real-time changes to production code create security risks and the potential for significant failures”.[14] The correct response is not to argue with that. It is to design the engagement so that the claim is false in your case, and to be able to show how.
14 Where the model fails
Twelve failure modes. Each has symptoms you can look for, a root cause that is usually structural rather than individual, and an owner — because a mitigation with no owner is a hope.
Snowflake deployments Structural
Symptoms. Upgrades are scheduled per customer. Nobody can say what version any given deployment is on without asking. A platform release note has to be rewritten for each account.
Root cause. No extension contract, so every requirement was met by modifying whatever was nearest. The core and the customization were never separated because separating them cost time in week three.
Engineering consequence. The product becomes N products. Upgrade cost grows with the number of customers rather than staying flat, which is the specific economic property a software business depends on.
Mitigation. A versioned extension surface; a rule that custom code cannot reach into core internals; continuous integration for every extension against the next core release, with deployment blocked on failure; a customization debt register with a disposition per entry.
Owner. Platform engineering owns the contract. Delivery leadership owns whether the rule is enforced under pressure.
Hero engineering Structural
Symptoms. One engineer is on every incident call for an account. Documentation is thin because the person who would write it is the person answering the questions. A week of leave becomes a risk event.
Root cause. Ownership was assigned to a person instead of a team, and the incentive structure rewarded responsiveness over transferability.
Engineering consequence. The deployment becomes unmaintainable by anyone else, which means it also becomes unsellable as a repeatable offering.
Mitigation. Pods with shared ownership and deliberate rotation. Runbooks written during the engagement rather than at the end. A handoff exercise where someone else operates the system for a period while the original engineer stays silent.
Owner. Delivery management, not the engineer.
Permanent privilege Security
Symptoms. A vendor account with production administrator rights and no expiry. Access reviews that return “still required” every cycle without evidence. Incident response that depends on that account existing.
Root cause. Implementation access was granted broadly to unblock delivery and nobody was accountable for narrowing it once delivery finished.
Engineering consequence. A vendor credential compromise becomes a customer production incident, and the customer’s own privileged access model has an exception in it that its assessors will find.
Mitigation. Just-in-time elevation through the customer’s workflow, scoped and time-boxed, with session recording and audit custody outside the vendor’s reach. Access should narrow as the deployment matures; if it has not narrowed in a year, that is a finding.
Owner. Customer security, with the vendor obliged to design for it from the first sprint.
Prototype permanence Delivery
Symptoms. The demo is in production. Credentials in configuration files. No tests around the part everyone depends on. A service whose name still contains “poc”.
Root cause. The prototype worked, the sponsor saw it work, and no gate existed between working and deployed.
Engineering consequence. Every subsequent change is high-risk because nobody knows what the system does under conditions the prototype never met.
Mitigation. Treat production readiness as a distinct gate with an explicit checklist and a named approver. See Production readiness. Build prototypes so they are cheap to throw away, and say out loud at the start that they will be.
Owner. Whoever signs the production readiness review. If nobody signs it, it is not a gate.
Scope gravity Delivery
Symptoms. The engineer is fixing problems in adjacent systems they did not build. Requests arrive directly rather than through any intake. The original success measure has not been discussed for two months.
Root cause. An engineer who is present, capable and already inside the estate attracts work. Nothing about this is the customer behaving badly.
Engineering consequence. The engagement outcome is never reached, because effort went to a series of smaller outcomes nobody is measuring.
Mitigation. Written scope with explicit exclusions. An intake path for new requests. A standing agenda item comparing effort against the agreed outcome. Willingness to decline, backed by delivery management rather than left to the individual.
Owner. Delivery management and the customer sponsor jointly.
Product feedback black hole Structural
Symptoms. The same workaround appears in three deployments. Field engineers stop filing issues because nothing happens to them. Product roadmaps contain nothing traceable to a deployment.
Root cause. No route from field observation to roadmap that has a named owner and a response commitment on the product side.
Engineering consequence. This is the failure that turns the whole model into consulting. Marginal deployment cost never falls because nothing learned is ever built once.
Mitigation. A field findings log with a triage commitment. A metric on the product side counting roadmap items that originated in the field. Rotation of engineers between field and product in both directions.
Owner. Product leadership. Delivery can file findings; only product can act on them.
Customization debt Structural
Symptoms. Local requirements accumulate faster than reusable abstractions. A growing share of engineering time goes to maintaining things built for one customer. Nobody can produce a list of what exists.
Root cause. Customization is easy to create and nobody is accountable for retiring it. There is no register, so there is no visibility, so there is no decision.
Engineering consequence. Capacity is consumed by maintenance that was never planned or funded, and it is invisible in delivery reporting because it is booked as delivery.
Mitigation. The register described in Productization economics, reviewed on a cadence, with an explicit retirement decision per entry and time allocated to act on it.
Owner. Platform engineering, with delivery contributing entries.
IP boundary failure Commercial
Symptoms. Nobody can say cleanly which code is the vendor’s reusable intellectual property and which encodes the customer’s process. Contract renewal surfaces a question nobody has an answer to.
Root cause. Repository structure, licensing and contract language were not designed to keep the two separable, and the separation is much harder to establish retroactively than to maintain.
Engineering consequence. Reuse becomes legally risky in one direction and customer confidentiality becomes risky in the other. Both chill exactly the behaviour the model depends on.
Mitigation. Separate repositories with an explicit interface between them. Contract language covering ownership of work product, derived improvements and the vendor’s right to generalize a pattern. Review at engagement start, not at renewal.
Owner. Legal and product jointly, informed by architecture.
Compliance by architecture Governance
Symptoms. A slide claiming the design is compliant with a named regime. Technical controls presented as evidence with no assessment behind them. Nobody can produce the control narrative an assessor would ask for.
Root cause. Confusion between implementing a control, operating it, evidencing it, and having it assessed. These are four different things done by four different sets of people.
Engineering consequence. A gap discovered during assessment rather than during design, at the point where remediation is most expensive and most public.
Mitigation. The vocabulary in Architecture versus compliance. Architecture supports controls; organizations operate them; evidence demonstrates them; assessors assess them; authorizing officials accept residual risk.
Owner. Governance, risk and compliance, with architecture supplying mechanisms and evidence.
AI without evaluation AI
Symptoms. Quality is assessed by demonstration. No labelled evaluation set exists. A model version change ships without a regression run because there is nothing to run.
Root cause. Evaluation was treated as testing, and testing was scheduled after the thing worked rather than as the definition of working.
Engineering consequence. Nobody can tell whether the system got worse, so nobody can tell whether a change was safe. The model provider ships an update and the failure surfaces as user complaints.
Mitigation. A frozen evaluation set built from real cases before the prototype is accepted. Separate measurement of retrieval and generation. A regression gate on every model, prompt or retrieval change. Where an LLM is used as a judge, its agreement with human labels is itself measured — it “is not a silver bullet”.[73]
Owner. The engineer who owns the AI component, with the workflow owner supplying the labels.
Agent over-permission AI security
Symptoms. An agent holds a broad token because scoping was fiddly. Tools that can write, delete or transact are available in contexts that only needed to read. Nobody can enumerate the actions the agent can take.
Root cause. OWASP names the three roots directly: excessive functionality, excessive permissions, excessive autonomy.[43] Excessive Agency rose to third in the 2026 edition of its LLM Top 10, from sixth.[43]
Engineering consequence. A successful prompt injection reaches everything the token reaches, and the blast radius was set months earlier by a convenience decision nobody recorded.
Mitigation. Minimum tool set per context. Per-target token exchange with explicit audience rather than passthrough. Authorization enforced downstream, not by the model. Transaction and rate limits. Approval bound cryptographically to the specific payload rather than to the session.
Owner. Security architecture, enforced in the platform rather than in a guideline.
Engineer burnout People
Symptoms. Sustained after-hours incident load concentrated on the same names. Travel that never settles. Attrition that takes deployment knowledge with it.
Root cause. The model structurally concentrates context, customer pressure and operational responsibility in one place, and organizations under-invest in the mechanisms that spread it because those mechanisms slow delivery in the short term.
Engineering consequence. Mission continuity risk. When the person leaves, so does the only complete model of how the deployment works.
Mitigation. Rotation with real handover. On-call that includes people outside the engagement. Capacity planning that treats customer count per engineer as a managed number. Instrumented team-health metrics reviewed in the same meeting as delivery metrics.
Owner. Engineering leadership. This one cannot be delegated downward, because the people best placed to notice it are the people affected by it.
15 Reference scenarios
Composite architecture scenarios. None of the four below describes a client, a deployment or a delivered system. Each is assembled from constraints that recur across an industry, and each exists to show a trade-off rather than a result. There are no named organizations, no counts, no timelines, no savings and no outcomes, because I have none to report and inventing them is how this genre of writing usually goes wrong.
A Healthcare operations
Problem shape. Operational visibility requires combining a clinical record system, scheduling and device telemetry, while access is tightly constrained by role and the data carries the strictest handling obligations in the estate.
Architecture shape. Integration performed locally rather than by exporting records; standards-based clinical interoperability rather than direct database access; an operational semantic model separate from the clinical record; de-identification where the use case genuinely does not need identity; role-aware application surfaces; AI assistance confined to summarizing and surfacing rather than deciding.
The trade-off. Operational freshness against integration complexity and the width of the privacy boundary. Every increase in timeliness pulls more identifiable data closer to more people, and the honest version of this design states which of those it chose and why.
What I will not claim. That such a design is HIPAA compliant. Applicability depends on the system boundary, the covered entity’s own controls, the business associate agreement and an assessment nobody has performed.
B Cross-region financial operations
Problem shape. Analysts need a global operational view while source records remain subject to jurisdictional constraints and internal need-to-know rules that differ by region.
Architecture shape. Regional processing with the query pushed to the data; policy-aware federation returning aggregates rather than records; lineage preserved so an aggregate can be explained; identity-aware drill-down that resolves to detail only for a principal authorized in that jurisdiction.
The trade-off. Central analytical convenience against residency, privacy and internal access boundaries. Federated queries are slower, harder to optimize and harder to debug, and each regional node is a component someone has to patch.
What I will not claim. A regulatory outcome. The source material this project started from asserted an audit passed with zero findings; there is no regulator, scope, period or assessor attached to that claim and it is not repeated here.
C Remote industrial operations
Problem shape. High-volume telemetry is generated where connectivity is intermittent and expensive, and operational decisions cannot wait for a round trip to a cloud region.
Architecture shape. Inference local to the site; buffering that survives an outage measured in days; compressed telemetry prioritized by operational value rather than by timestamp; model deployment as a controlled, signed, versioned artifact; drift detection running locally with retraining requests queued for the next window.
The trade-off. Centralized compute efficiency against local autonomy and deterministic behaviour. The local model will be smaller and less capable than the one you could run centrally, and that is the price of it working when the link is down.
What I will not claim. Prevented failures or financial savings. Prevented incidents are counterfactual by construction; a saving derived from them is an estimate dressed as a measurement.
D Disconnected, high-assurance environment
Problem shape. Software must operate with restricted or absent connectivity, under tightly controlled data movement, where remote vendor access is not available at all.
Architecture shape. A fully local runtime with no dependency on a remote control plane; deployment artifacts signed and verified before load; synchronization through a controlled, reviewed path rather than an open channel; policy decisions made locally because the policy service cannot be reached; observability that collects locally and is exported deliberately.
The trade-off. Centralized management against local sovereignty and assurance. Everything that a control plane would have done for you becomes a local procedure that a local operator has to execute correctly.
What I will not claim. Any classified deployment, agency or network. The source material named a specific classified network in an unattributed narrative; that claim is not repeated, and neither is the security clearance requirement it stated as universal, which is contract- and nation-specific.
16 When should an organization use forward deployment?
Up to nine questions, depending on the path. The honest answer for most problems is one of the six alternatives, and an organization that reaches “forward deployed engineering” for everything has stopped using the question as a filter.
The alternatives are not lesser outcomes. Everest Group’s framing is that the question is “not whether they are good or bad” but “where they fit”, and that in environments with mature change control forward deployed engineers “are not only unnecessary… but they can be actively harmful”.[14] Choosing partner delivery for a repeatable implementation is a better decision than assigning the scarcest engineers you have to work a certified third party could do without them.
17 Operating-model maturity
Six levels. Most organizations running this model are at level 0 or 1 and describe themselves as being at level 3, which is easy to test: the artifacts of the higher level do not exist and can be asked for.
Individual engineers, undocumented access, manual deployment
What it looks like. Deployment knowledge lives in people. Access was granted informally and has not been reviewed. Releases happen when someone runs something from a laptop. Customer-specific code exists in places nobody has catalogued.
Primary risk. Dependency on people. A departure takes the deployment’s only complete mental model with it.
Next action. Get everything into version control, write down what access exists, and produce one runbook that someone else can follow end to end.
Engagement templates, source control, basic pipelines, defined access
What it looks like. Engagements start from a template. Code is in git with review. A pipeline builds and deploys. Access is defined even if it is broader than it should be. Runbooks exist and are occasionally wrong.
Primary goal. Make a successful deployment repeatable rather than remarkable.
Next action. Introduce a production readiness gate with a named approver, and start the customization register before you need it.
Formal intake, architecture and security gates, evaluation, SLOs
What it looks like. Work enters through an intake with qualification. Architecture and security review happen before build rather than after. AI components have evaluation sets. Services have service level objectives someone watches. Reusable components exist and are used.
Primary goal. Control risk without destroying delivery speed. This is the level where the trade-off is real and where governance most often overshoots.
Next action. Measure how long the gates take. A gate that adds weeks and catches nothing gets routed around, and the routing around is invisible until something fails.
SDKs, golden paths, policy as code, deployment automation, feedback pipeline
What it looks like. There is a supported way to build the common thing and it is faster than the unsupported way. Policy is expressed as code and evaluated in the pipeline. Deployment is automated across topologies. Field findings reach product through a route with a response commitment. An extension catalogue exists with owners.
Primary goal. Reduce the marginal effort of the next deployment.
Next action. Measure marginal effort. If it is not falling, the platform investment is not yet doing what it was funded to do.
Partners and customer engineers handle common deployments
What it looks like. Certified partners or the customer’s own engineers deliver the well-understood patterns. The scarce internal engineers are deployed against problems that are genuinely new. Productization is a funded activity with a backlog rather than a side effect. Engagement analytics are standardized enough to compare accounts.
Primary goal. Concentrate scarce talent where it creates new leverage instead of where it is most requested.
Next action. Check that the partner channel is delivering the same quality. This is the level where quality quietly diverges and nobody measures it.
Field signal shapes product priorities; deployment is a source of advantage
What it looks like. Deployment telemetry informs architecture decisions. Repeated field patterns become platform primitives on a predictable cadence. The economics of an engagement improve measurably from one to the next, and someone can show the series.
Primary goal. Make deployment itself a source of product advantage rather than a cost of sale.
Honest caveat. I have seen this described more often than demonstrated. The evidence that an organization is here is a chart of marginal deployment cost over time, and most organizations cannot produce that chart.
18 Who owns what
An example operating model, not a universal one. Every organization I have looked at places at least two of these differently, and the value of writing it down is the argument it starts rather than the grid it produces.
| Activity | FDE | Platform eng | Product eng | Research | Vendor security | GRC | Solutions arch | Cust. product owner | Cust. security | Cust. operations |
|---|---|---|---|---|---|---|---|---|---|---|
| Discovery | R | I | I | I | I | I | C | A | C | C |
| Scoping | R | C | I | I | C | C | C | A | C | C |
| Architecture | A | C | C | I | C | C | R | C | C | I |
| Integration | R | C | I | I | I | I | I | C | C | A |
| Application development | A | C | C | I | I | I | I | C | I | I |
| Model selection | R | C | C | A | C | C | I | C | C | I |
| Evaluation | A | C | C | C | C | C | I | R | I | C |
| Security review | C | C | I | I | R | C | C | I | A | I |
| Data authorization | C | I | I | I | C | C | I | C | A | R |
| Production deployment | R | C | I | I | I | I | I | C | C | A |
| Incident response | C | C | I | I | C | I | I | I | C | A |
| SLO ownership | C | R | C | I | I | I | I | C | I | A |
| Change management | R | C | I | I | C | C | I | C | C | A |
| Adoption | C | I | I | I | I | I | C | A | I | R |
| Productization | C | R | A | C | I | I | C | I | I | I |
| Handoff | R | C | I | I | I | C | I | A | C | C |
Three placements in this grid are deliberate and are where most disagreement lands. Production deployment and incident response are accountable to customer operations rather than to the vendor engineer, which is the structural expression of the access-boundary argument: the party accountable for production is the party that controls production. And evaluation is accountable to the engineer who built the AI component but responsible to the customer product owner, because the labels that define correct behaviour are the customer’s knowledge and cannot be delegated to whoever wrote the code.
19 Prototype is not production
A working prototype is evidence that a hypothesis held under the conditions you tested. It is not a production system, and the gap between them is a defined body of work rather than a tidying-up exercise. This is the checklist I would put a named approver behind.
Application 10 checks
- Automated tests covering the paths the workflow actually depends on, not the paths that were easy to test.
- Error handling that distinguishes retryable from terminal, and surfaces the difference to the caller.
- Explicit versioning of the deployed artifact, traceable to a commit.
- A rollback that has been executed at least once, in a real environment, by someone who was not the author.
- Configuration separated from code, with environment-specific values held outside the artifact.
- Startup behaviour defined when a dependency is unavailable — fail fast, degrade, or wait, decided rather than inherited.
- Resource limits set, and behaviour under limit understood.
- Idempotency for every operation that can be retried.
- No credentials, endpoints or customer identifiers in source.
- A named owner in the repository, not a team alias.
Data 7 checks
- Lineage from every field the workflow depends on back to its system of record.
- Freshness expectations stated per source, and monitored against.
- Quality checks that fail loudly rather than propagating nulls into a decision.
- Retention and deletion defined per dataset, and implemented rather than documented.
- Classification applied, with handling rules that follow the classification through transformation.
- Backfill and replay behaviour defined for the case where a source was wrong.
- A stated position on what happens to derived data when the source record is deleted.
AI components 8 checks
- An evaluation set built from real cases, frozen, versioned, and owned by someone who understands the workflow.
- A recorded baseline for every quality dimension you intend to defend.
- Regression gates that run on model change, prompt change and retrieval change — all three, independently.
- Groundedness measured separately from task success, because retrieval failure and generation failure need different fixes.[62]
- Tool-use tests that assert what the system does not call as well as what it does.
- Safety and refusal behaviour tested against adversarial inputs, including content that reaches the model through retrieval.
- An abstention path that is evaluated as a correct outcome rather than a failure.
- Model version, prompt version and retrieval configuration recorded on every request, or the telemetry is not diagnostic.
Identity 5 checks
- Distinct service identities per component, with no shared credential across environments.
- Human roles defined against duties rather than against convenience.
- Least privilege verified by testing that an over-broad action fails, not by reading the policy.
- An access lifecycle with joiners, movers, leavers and a review that can result in removal.
- No standing vendor access to production, and a documented just-in-time path that has been exercised.
Security 7 checks
- A threat model that names the assets, the boundaries and the assumptions, reviewed with the customer’s security team.
- Dependency scanning in the pipeline, with a policy for what blocks a release.
- Secrets held in a managed store with rotation that has been performed rather than scheduled.
- Encryption in transit and at rest, with key custody separated from platform administration.
- Security logging that reaches the customer’s own detection capability, in a format it can consume.
- Integration with the customer’s incident process, including who is called and what they are authorized to do.
- Signed build artifacts with provenance, and a verification step that actually rejects an unsigned one.
Reliability 6 checks
- A service level objective that reflects what the workflow needs, agreed with the person who owns the workflow.
- Capacity understood at peak, not at average.
- Failure modes enumerated per dependency, with the chosen behaviour for each.
- Retry strategy with backoff and a bound, so a downstream outage does not become a self-inflicted denial of service.
- Backup and restore proven by restoring, in a drill, with the time recorded.
- A disaster recovery position that names what is lost and how much time it takes, rather than asserting that nothing is.
Operations 5 checks
- Dashboards that answer “is it working” before they answer “what is it doing”.
- Alerts tied to symptoms users would notice, with a documented action for each.
- Runbooks written by someone and executed by someone else at least once.
- An escalation path that reaches a person, with hours and expectations stated.
- Named operational ownership on the customer side, agreed before go-live rather than discovered during the first incident.
Governance 5 checks
- An approved-use statement for the system, including what it is explicitly not for.
- A recorded data-use decision covering purpose, lawful basis where applicable, and any secondary use.
- An AI risk decision recorded against whatever framework the organization uses, with the residual risk accepted by a named person.
- Release approval with an accountable approver rather than a group.
- Evidence retention defined, so the record still exists when it is asked for.
Cost 5 checks
- Infrastructure cost attributed to the workload rather than to a shared account.
- Model and token spend measured per successful workflow, not per call.
- Observability cost budgeted, because high-cardinality AI telemetry is expensive and gets cut first when nobody planned for it.
- Data movement cost understood, especially in federated and hybrid designs.
- Ongoing support burden estimated and staffed, or the engagement never ends.
Adoption 5 checks
- A named workflow owner on the customer side who wants this to exist.
- Training that reaches the people who will use it, not only the people who sponsored it.
- A user experience tested with users rather than with the project team.
- A success metric agreed in advance, measurable without a special report.
- A feedback mechanism that produces work rather than sentiment.
20 Measuring outcomes without rewarding heroics
No benchmark values appear below. Every organization’s baseline differs and a published target would be a number I invented. What each metric is for is stated instead, because a metric whose purpose is unstated gets optimized in whatever direction is easiest.
Delivery
- Time to validated prototype
- Reveals whether discovery is converging or circling. A long time here is usually a scoping problem, not an engineering one.
- Time to production
- Reveals the size of the readiness gap. If it dwarfs time to prototype, the prototype was not built toward production.
- Blocker age
- Reveals dependency on the customer organization. Old blockers are an escalation signal, not an engineering signal.
Adoption
- Active users in the target population
- Reveals whether the workflow owner’s people actually adopted it, as opposed to whether accounts were provisioned.
- Workflow completion rate
- Reveals whether the system finishes the job or hands it back part-done, which is the failure users stop reporting.
- Repeat usage
- Reveals whether it earned a place in the routine. First use is curiosity. Repeat use is the signal.
Reliability
- SLO attainment
- Reveals whether the target was set honestly. Permanent 100% attainment usually means the objective is not binding.
- Incident rate and MTTR
- Reveals operational maturity. Watch the trend rather than the level.
- Change failure rate
- Reveals whether the readiness gate is doing anything.
AI quality
- Evaluation pass rate
- Reveals regression on a frozen set. Only meaningful if the set has not been edited to make it pass.
- Groundedness
- Reveals whether answers are supported by retrieved evidence, measured separately from whether they were useful.
- Task success
- Reveals workflow effectiveness, which benchmark scores do not predict.[69]
- Human escalation rate
- Reveals where the system stops being autonomous. A falling rate with flat quality is the good direction; a falling rate with unmeasured quality is a warning.
Economics
- Cost per successful workflow
- Reveals the real unit economics. Cost per call flatters systems that fail cheaply and retry often.
- Model and token cost
- Reveals where routing or caching would pay. Routing is an established pattern with published trade-offs, and the savings are highly workload-dependent.[72]
- Engineering effort per deployment
- Reveals whether the platform investment is working. This is the single number that separates a product from a consultancy.
Productization
- Reusable component ratio
- Reveals how much of a deployment came from the shelf.
- Duplicate customizations
- Reveals patterns that should have been promoted and were not. The third occurrence is the signal.
- Custom code retired
- Reveals whether the debt register is a decision-making tool or a list.
- Field patterns converted to platform capability
- Reveals whether the loop in Figure 3 is closed.
Product feedback
- Actionable field findings
- Reveals whether engineers are still filing, which they stop doing when nothing happens.
- Roadmap items originating in the field
- Reveals whether product is receiving. This is the counterpart metric and both are needed.
- Regressions found through deployment
- Reveals the value of the field as a test surface, which is real and rarely counted.
Team health
- Customer load per engineer
- Reveals concentration before it becomes attrition.
- After-hours incident load
- Reveals whether operational ownership actually transferred at handoff or only formally.
- Travel burden
- Reveals sustainability. Worth tracking even where embedding is mostly remote.
- Ownership concentration
- Reveals bus factor per deployment. Should be reviewed in the same meeting as delivery metrics, not in a separate one nobody attends.
21 Standards and governance alignment
This is a crosswalk, not a conformity claim. Nothing in this architecture makes an organization compliant with anything. Applicability depends on the system boundary, the organization’s own controls, contractual obligations, the regulatory context and an assessment that has not been performed. The columns are deliberately named “architecture mechanism”, “operational process” and “example evidence” to keep those three things apart.
Versions below were checked against primary sources on 9 August 2026. Several changed within the preceding twelve months in ways that invalidate widely circulated guidance, and those are marked. Filter to the frameworks that apply to your boundary rather than reading all of them.
| Framework | Concern it speaks to | Architecture mechanism | Operational process | Example evidence | Caveat |
|---|---|---|---|---|---|
| NIST AI RMF 1.0 AI 100-1, Jan 2023 | Whether an AI use case has been governed, mapped, measured and managed rather than merely built | Model inventory; use-case classification; evaluation harness; human oversight points; rollback to a prior model or a non-AI path | Intake and risk classification; recorded acceptance of residual risk by a named person; periodic re-review on model change | Risk decision record; evaluation results per release; oversight design rationale | Final, but NIST states the framework is being revised, with no public draft as of August 2026.[27] The revision was directed by US federal AI policy in July 2025.[28] Do not cite a revision that has not been issued. |
| NIST AI 600-1 GenAI Profile, Jul 2024 | Generative-specific risks including confabulation, data leakage and information integrity | Grounding and provenance; output validation; egress control on generated content; abstention path | Pre-deployment red teaming; content-integrity review for consequential decisions | Adversarial test results; groundedness measurements; incident records | NIST uses “confabulation” rather than “hallucination” and flags the risk as sharpest in consequential decision-making.[29] A companion profile to the AI RMF, so its standing follows the revision. |
| NIST CSF 2.0 CSWP 29, Feb 2024 | Whether the engagement has an owner, and whether outcomes are governed rather than assumed | The whole architecture read as Govern, Identify, Protect, Detect, Respond and Recover surfaces — Govern is the function added in 2.0 and the one an engagement is most likely to leave unassigned | Governance function assignment; third-party risk treatment for the vendor engineering team | Current and target profiles; supplier risk assessment | CSF 2.0 says explicitly that it “does not prescribe how outcomes should be achieved”.[23] It is a structure for the conversation, not a control set. |
| NIST SP 800-53 Rev. 5, Release 5.2.0 | The control vocabulary an assessor will use for access, audit, configuration and supply chain | AC, AU, CM, IA, SC and SR family implementations across the stack | Control selection against a baseline; implementation; assessment; continuous monitoring | Control implementation statements; assessment results; monitoring output | Cite the release, not just the revision. Release 5.2.0 (August 2025) added SA-15(13), SA-24 and SI-02(07) in response to EO 14306.[24] |
| NIST SP 800-207 Zero Trust, Aug 2020 | Per-request access decisions where the network is assumed compromised | Policy decision and enforcement points; workload identity; per-session authorization; continuous verification | Trust algorithm definition; policy authoring separated from policy administration | Policy definitions; authorization decision logs; asset posture data | “ZT is not a single architecture but a set of guiding principles.”[21] NCCoE SP 1800-35 became final in June 2025 and consolidated the earlier draft volumes.[22] |
| ISO/IEC 27001 2022, incl. Amd 1:2024 | Whether the provider’s own delivery process, including field engineering, sits inside a managed system | Technical controls in Annex A relevant to access, cryptography, operations and supplier relationships | ISMS scope covering forward deployed operations; risk treatment; internal audit; management review | Statement of Applicability; risk treatment plan; audit records | Cite as ISO/IEC 27001:2022 including Amendment 1:2024. The transition deadline passed in October 2025, so a certificate against the 2013 edition is no longer valid.[31] |
| ISO/IEC 42001 AI management systems, 2023 | Whether AI is managed as a system with objectives, roles and improvement rather than per project | AI system inventory; impact assessment inputs; lifecycle controls | AI policy; objectives; competence; operational planning; internal audit | AI management system documentation; impact assessments; audit records | ISO/IEC 42006:2025 sets requirements for bodies certifying against 42001, which is what makes accredited certification operational.[32] Certification is of the management system, not of a model or an architecture. |
| OWASP Top 10 for LLM Applications 2026 edition — renumbered | The application-level risks an AI deployment is most likely to carry into production | Downstream authorization; output handling; retrieval permission checks; consumption limits | Threat modelling per use case; security testing including retrieval-borne injection | Test results; findings tracked to closure | Changed in August 2026. Identifiers moved from LLM##:2025 to LLM##:2026; Excessive Agency rose to third; a new entry covers hidden context exposure.[43] Any matrix citing 2025 identifiers is stale. |
| OWASP Top 10 for Agentic Applications ASI01–ASI10, Dec 2025 | Risks specific to systems that plan, remember and act — goal hijack, tool misuse, privilege abuse, cascading failure | Tool permission scoping; memory and context integrity; inter-agent authentication; action validation and limits | Agent registration and approval; permission review; kill-switch procedure | Tool permission matrix; action audit trail; incident records | A distinct series from the LLM list, with its own identifiers. It did not exist a year ago.[44] |
| MITRE ATLAS data release v2026.05 | Adversary tactics and techniques against AI systems, modelled as ATT&CK is for enterprise | Detection coverage mapped to technique identifiers; telemetry sufficient to see them | Threat-informed defence; detection engineering; purple-team exercises | Technique coverage map; detection test results | Continuously updated rather than versioned like a standard; techniques are now tagged by platform including agentic AI.[46] Cite the data release and the date. |
| CIS Critical Security Controls v8.1, Jun 2024 | A prioritized baseline for the operational hygiene an engagement inherits | Asset and software inventory; secure configuration; account and access management; audit log management | Implementation group selection; safeguard implementation and review | Inventory records; configuration baselines; log management evidence | v8.1 added a governance function and realigned mappings to CSF 2.0.[34] Stable since June 2024, which not every row here can say. |
| SLSA v1.2 — v1.1 retired | Whether the artifact deployed into a customer environment is the one that was built from the reviewed source | Provenance generation; artifact signing; verification at deployment that rejects unsigned or unattested builds | Build platform hardening; provenance policy; exception handling | Provenance attestations; verification logs; policy definition | v1.1 is explicitly retired. v1.2 was approved in November 2025 and adds the source track alongside the build track.[35] |
| OpenSSF OSPS Baseline release 2026-02-19 | Baseline security practice for the open-source components an engagement pulls in | Dependency inventory and SBOM generation; component provenance; vulnerability scanning in the pipeline | Component intake policy; upgrade cadence; end-of-life handling | SBOMs; scan results; component approval records | Date-versioned rather than semantically versioned; the February 2026 release added controls and expanded external mappings.[36] |
| SOC 2 TSC 2017, points of focus rev. 2022 | How a customer gains assurance over controls at a vendor that has access to its systems | Logical access controls; change management; monitoring; the access-broker design in Figure 7 | Control operation over a defined period; management’s description and assertion | The examination report itself; control operation evidence for the period | SOC 2 is an examination, not a certification. There is no such thing as being “SOC 2 certified”, and it says nothing about whether an architecture is sound.[33] Security is the mandatory category. |
| HIPAA Security Rule 45 CFR 164 Subpart C | Only where electronic protected health information is inside the boundary | Access control, audit controls, integrity, transmission security; de-identification where the use case permits | Risk analysis; workforce and business associate management; sanction and review policies | Risk analysis record; BAA; access review evidence | The proposed 2025 overhaul has not been finalized and has moved to the long-term actions list; the existing Security Rule governs.[37] Guidance predicting a 2026 final rule is now wrong. |
| PCI DSS v4.0.1 | Only where cardholder data enters the system boundary | Segmentation that removes systems from scope; tokenization; MFA for all access to the cardholder data environment | Scope confirmation; targeted risk analyses; authenticated internal scanning | Scope documentation; scan results; the assessment itself | The previously future-dated v4.x requirements have been in force since 31 March 2025.[38] Describing them as upcoming is a year out of date. |
| FedRAMP 20x, Phase 3 active | Only for cloud offerings serving US federal agencies | Cloud-native architecture, identity, logging and recovery expressed as continuously validated indicators | Continuous validation rather than point-in-time authorization packages | Machine-readable evidence against Key Security Indicators | Terminology changed. “Authorization” is now Certification, impact levels are Certification Classes A–D, the SSP is a Security Decision Record and POA&M items are Accepted Weaknesses.[39] There is no mechanism by which an architecture is pre-accredited. |
| CMMC and NIST SP 800-171 Phase II suspended Jul 2026 | Only where controlled unclassified information is in scope for a defence supply chain | The 800-171 requirement families as implemented in the deployment | Self-assessment, scoring and affirmation under the phase currently in force | Assessment score; system security plan; affirmation record | CMMC Phase II was suspended on 13 July 2026; only Phase 1 self-assessment is in force.[40] Note also that NIST’s current publication is SP 800-171 Rev. 3 while CMMC Level 2 is still assessed against Rev. 2.[41] |
22 Architecture versus compliance, worked once
The distinction the previous section rests on is easier to see in a single example than in an argument. Privileged access is the right one to use, because it is the control a forward deployed engagement is most likely to weaken and the one an assessor will look at first.
| Layer | Privileged access, worked through |
|---|---|
| Requirement area | Vendor engineers need occasional access to a customer production environment to diagnose and repair. |
| Architecture mechanism | Federated identity from the customer’s provider with multi-factor authentication; a just-in-time role that does not exist until requested; privileged access management brokering the session; session recording; audit records written to a store the vendor cannot administer. |
| Operational process | An access request raised in the customer’s own workflow with a stated reason and duration; approval by someone other than the requester; automatic expiry; periodic review that can and sometimes does result in removal. |
| Evidence | The access request and its approval; the authorization decision log; the role definition showing scope; the session recording; the access review record showing what was removed. |
| Assurance | Someone independent tests whether the control operated as designed over a period — that expiry actually expired, that reviews actually removed access, that the recording exists for a sampled session. |
| Authorization | An accountable person accepts the residual risk. That decision is theirs and cannot be produced by a diagram. |
Each row is a different activity performed by different people. The architecture supplies the second row and makes the fourth row possible. It cannot supply the third, the fifth or the sixth. NIST puts the responsibility explicitly on the organization: “organizations have the responsibility to select the appropriate security and privacy controls, to implement the controls correctly, and to demonstrate the effectiveness of the controls in satisfying security and privacy requirements”.[24] Assessment is defined as determining “if the controls are implemented correctly, operating as intended, and producing the desired outcomes”.[25] None of those verbs belongs to a design.
“Security and privacy control assessments are not about checklists, simple pass/fail results, or generating paperwork.”
NIST SP 800-53A Rev. 5[26]
23 What a disciplined engagement produces
Nineteen artifacts. None of them are downloadable from this page, because I have not written them for a real engagement and publishing templates I have not used would be filler. What follows is what each one is for and why its absence hurts.
01 Engagement charter
States the outcome, the boundary, the exclusions and who decides. The document people stop reading in month two and should reread in month four.
02 Discovery findings
What was observed rather than what was requested. The gap between the two is usually the whole value of the discovery phase.
03 Context diagram
The system, the actors and the neighbouring systems on one page. If it needs two pages, the boundary is wrong.
04 Architecture decision records
Decision, context, options considered, consequence. Their value is in year two, when someone asks why and the person who knew has left.
05 Integration inventory
Every system touched, its owner, its protocol, its availability and its data classification. The artifact that makes change impact assessable.
06 Data-flow diagram
Where data moves, crosses a boundary, is transformed and is retained. Prerequisite for both the threat model and the privacy analysis.
07 Threat model
Assets, boundaries, adversaries, assumptions. Its most useful section is the assumptions, because those are what change.
08 Identity and access matrix
Who and what can do what, in which environment. The document that makes over-permission visible instead of theoretical.
09 Data classification map
Classification per dataset and the handling that follows from it, tracked through transformations rather than stated at the source.
10 Evaluation plan
For AI components: what is measured, on which frozen set, by whom, and what result blocks a release.
11 Production readiness review
The checklist in section 19 with a name against it. If nobody signs it, it has not gated anything.
12 SLO definition
What the workflow needs, agreed with the person who owns the workflow, with the error budget and what happens when it is spent.
13 Risk register
Open risks with owners and treatment. Distinct from the issue log; a risk that has occurred is an issue and should move.
14 Cutover plan
Sequence, dependencies, verification steps, rollback trigger and the person authorized to pull it. Written before the day.
15 Operations runbook
Written by one person and executed by another before go-live. A runbook nobody has followed has not been tested.
16 Incident escalation matrix
Who is called, in what order, with what authority, and what the vendor is permitted to do without waiting.
17 Customization debt register
The register from section 09. The artifact that makes upgrade planning a calculation rather than an archaeology exercise.
18 Product feedback log
Field findings with a triage commitment attached. Without the commitment, the log stops receiving entries, and the engineers who stopped filing were right to.
19 Handoff package
Everything above, plus the demonstration that someone else operated the system. Handoff has to be proven by someone else running the system, which is why it belongs in this list rather than in a folder.
24 What the model teaches
Eight conclusions I would defend, written as claims rather than as principles because principles are harder to disagree with and disagreement is the useful part.
- Proximity to the customer is not a substitute for engineering discipline, and is often used as one. The argument that tests, reviews and documentation can wait because the customer needs it Friday is the same argument every quarter, and it compounds.
- The hardest problems sit between systems and between teams, not inside an API. That is why the role exists at all, and why staffing it with someone who cannot change code produces a well-documented list of things that cannot be done.
- A successful prototype is evidence that a hypothesis held, not evidence that a system exists. Treating the two as the same thing is how a project comes to overrun its estimate by a multiple rather than a margin.
- The model scales only when repeated fieldwork becomes reusable platform capability. Everything else in the operating model is downstream of whether that loop is closed and instrumented. Nobody has ever closed it by intending to.
- Access should narrow as a deployment matures. If the vendor account still holds the permissions it needed during implementation, the engagement has a security posture that was set by delivery pressure two years ago.
- AI deployment requires continuous evaluation because model behaviour is part of the production system. A component whose vendor can change its behaviour without your release cycle is a dependency with no change control, and evaluation is the only thing standing in for it.
- A semantic layer improves retrieval, provenance and policy enforcement, and does not eliminate probabilistic failure. The measurements are unambiguous and they are worse than the diagrams imply. Design for the wrong-but-plausible answer, because it will occur and it will look correct.
- The strongest organizations optimize two things at once: the customer outcome, and the rate at which field learning improves the product. Optimize only the first and you have built an excellent consultancy. Optimize only the second and you have built a platform nobody has deployed.
25 Questions
The role
What exactly is a forward deployed engineer?
A software engineer who works inside a customer’s operational context with the authority to design and build production systems there, and who stays accountable for how those systems behave. The distinguishing variables are depth of technical intervention and persistence of presence, not job title or location.
Is it just another name for consulting?
Sometimes, and the industry does not agree with itself. Traditional consulting typically produces recommendations that a separate team implements; the forward deployed model has the same people design, build and operate. But Databricks files its team under professional services,[8] Vercel’s role reports to a Director of Professional Services while the posting says it is not consulting,[9] and OpenAI and Anthropic both moved delivery into separate services ventures.[6][7] The useful question is not what it is called but whether the engineers have production authority and whether their learning reaches the product.
How is it different from a solutions architect?
A solutions architect is accountable for a design; a forward deployed engineer is accountable for a running system. The architect’s deliverable is a decision; the engineer’s is behaviour in production. Both roles exist for good reasons and the failure mode is asking one to do the other’s job with the other’s authority.
Does the role have to be on-site?
No. Embedding is operational, not geographic. Palantir’s own posting for the role caps travel at 25%,[2] Anthropic’s federal posting states 25–50%,[4] and PostHog runs its forward deployed engineer fully remote and asynchronous by deliberate choice.[12] The one topology where physical presence is genuinely forced is the disconnected environment, and even there the person present may be a cleared local operator rather than the vendor’s engineer.
How much coding is actually involved?
It varies enough that any single figure would be misleading. One independent reading of a major cloud vendor’s posting estimated roughly a quarter coding, half integration work, and a quarter meetings.[13] The number that matters is not the percentage but whether the engineer has authority to merge to production. Without that, the role is advisory regardless of how much code gets written.
What skills actually distinguish an effective one?
Software engineering depth is necessary and not sufficient. The differentiators are the ability to scope an ambiguous problem into something buildable, comfort operating inside another organization’s politics without becoming part of them, willingness to write the boring artifacts, and the judgment to distinguish a requirement that should be built from a request that should be declined.
Can a solutions architect move into the role, or an engineer move back into product?
Both happen and both are useful. Architects moving in need to rebuild the habit of shipping and being on call for what they shipped. Engineers moving to product carry a specific advantage, which is that they have watched the product fail in ways the roadmap did not anticipate. Deliberate rotation in both directions is one of the few structural mitigations for the feedback-loop failure that actually works.
Scope and economics
What makes a problem suitable for this model?
Strategic importance, genuine workflow ambiguity, integration complexity that standard configuration cannot reach, a production outcome that requires software to be written, and a reasonable chance the work reveals a reusable platform capability. Missing the last one is acceptable occasionally and fatal as a pattern.
When should professional services or a partner do it instead?
When the pattern is known and the work is implementation rather than discovery. A repeatable deployment handed to a certified third party is a better use of everyone than assigning your scarcest engineers to it. Everest Group goes further and argues that in environments with mature change control, forward deployed engineers can be “actively harmful” because real-time production changes bypass the safeguards those environments depend on.[14]
Is the model economically viable at scale?
Not as a universal delivery mechanism, and the people who run it say so. The viable shape is selective: forward deployment for landings, frontier problems and reference implementations, with repeatable patterns moving to a lower-cost tier or a partner. F-Prime Capital offers a threshold worth arguing about — if more than 30–40% of deployments need significant forward deployed effort, the problem is product design rather than go-to-market.[15] Treat that as a prompt to measure, not a target.
How do you stop customer-specific code becoming permanent debt?
Register it at creation with an owner, a core API dependency and a target disposition; run continuous integration for every extension against the next core release and block deployment on failure; review the register on a cadence with time allocated to act on it. The mechanism is not complicated. What fails is that nobody is accountable for retirement, so every entry defaults to permanent.
Who owns code produced during an engagement, and how should IP boundaries work?
Decide at engagement start and structure the repositories to match. Vendor reusable intellectual property and customer-specific process logic belong in separate repositories with an explicit interface, and the contract needs to cover work product ownership, derived improvements and the vendor’s right to generalize a pattern. Establishing this retroactively is materially harder than maintaining it, and the moment it is discovered to be missing is usually a renewal negotiation.
Security and access
Should engineers hold production administrator privileges?
Not as a standing arrangement. Just-in-time elevation requested through the customer’s own workflow, scoped to a role, time-boxed, session-recorded, with audit custody outside the vendor’s administrative reach. The cost is minutes at the start of an incident. The benefit is that a vendor credential compromise stops being a customer production incident.
How should access work in regulated environments?
The same way, with more evidence. Co-design the roles with the customer’s security and compliance teams rather than requesting access and negotiating afterwards. Expect to be managed as a third party or sub-processor, expect audit rights in the contract, and expect a joint governance forum reviewing architecture and changes. The engagements that go badly are the ones where security was engaged after the design was fixed.
How does the model align with DevSecOps and Zero Trust?
DevSecOps is the delivery discipline that keeps the model from producing snowflakes: pipeline-enforced testing, scanning, policy as code, signed artifacts and provenance. Zero Trust is the access posture that keeps it from producing standing privilege. Neither is optional and neither is produced by the architecture on its own — see section 12 for what a data-layer policy engine does and does not cover.
How does tool integration through MCP change the security architecture?
It converts an integration layer into something closer to a control plane, with a non-deterministic caller. The specification is explicit that tool descriptions are untrusted unless the server is trusted, and that the protocol cannot enforce its own security principles.[50] Practical minimum: an inventory of trusted servers with named approvers, per-target token exchange with explicit audience rather than passthrough, tool permissions scoped to the call, action validation, transaction limits, and audit that survives compromise of the audited system.
What should happen before an agent takes a real-world action?
Authorization enforced in the downstream system against the acting principal, not a decision made by the model. Validation of the action’s structure rather than its apparent intent. A limit on value and rate. Human approval where the action is irreversible or high-value, bound to that specific payload rather than to the session. And a log entry that records the model version, prompt version, retrieved context, tool call and authorization decision, because without those the incident review has nothing to work with.
AI systems
How is AI forward deployment different from normal application implementation?
Four differences. Quality is statistical rather than binary, so acceptance needs an evaluation set instead of a test suite. Behaviour changes when the model changes, on the vendor’s schedule rather than yours. Cost varies per request. And the system reads untrusted content, which means content is an attack surface. Everything else is a normal, difficult integration project.
Does retrieval eliminate hallucination?
No. One annotated corpus found 43.1% of responses in retrieval-augmented settings contained at least one hallucination.[60] Even when the source document is supplied directly in the prompt, the best measured groundedness scores sit in the low-to-mid eighties.[61] Retrieval substantially reduces the rate. It does not remove the failure mode.
Does a graph or ontology eliminate hallucination?
No, and the vendor most associated with the claim does not make it. Palantir’s own material is titled Reducing Hallucinations and pairs the ontology with human oversight.[53] Peer-reviewed work on knowledge graphs and LLMs frames the contribution as mitigation.[67] There is also a formal argument that hallucination cannot be eliminated in general,[63] and a competing position from OpenAI that confident hallucination is reducible because a model can abstain.[64] Both agree accuracy never reaches 100%.
What role does an enterprise ontology actually play, and is one required?
It gives you shared semantics across fragmented sources, a place to express authorization over relationships, and provenance that points at identified records. Palantir describes its Ontology as an operational layer connecting integrated digital assets to their real-world counterparts.[52] It is not required. It is justified when questions are relationship-heavy, entities are resolvable, and somebody will own the schema. Build one because you need those properties, not because the architecture diagram has a slot for it.
How should evaluation be handled?
A frozen set built from real cases before the prototype is accepted; retrieval and generation measured separately; regression gates on model, prompt and retrieval changes independently; abstention scored as a correct outcome. Where an LLM acts as judge, measure its agreement with human labels first — the practitioner consensus is that it “is not a silver bullet”.[73] Multi-dimensional evaluation is the established academic position, not a novelty.[71]
Why not just trust benchmark scores?
Because they measure benchmarks. A 2025 review found task-setup and reward-design flaws in widely used agentic benchmarks capable of shifting reported performance by up to 100% in relative terms, including one that scored empty responses as successes.[70] And the METR trial found real developers slowed down while believing they had sped up.[69] Evaluate against your own workflow.
Is the system deterministic if the query execution is?
One stage is. Given a fixed query and a fixed data state, execution is reproducible and auditable, which is genuinely valuable. The stages either side of it are language-model inference and are not deterministic — not even at temperature zero in practice, without batch-invariant serving.[74] Claim the determinism you have and no more.
Governance and organization
Does following NIST or ISO make the system compliant?
No. Architecture supports controls; organizations implement and operate them; evidence demonstrates them; assessors assess them; an accountable person accepts residual risk. NIST places selection, correct implementation and demonstration of effectiveness on the organization.[24] A design cannot perform any of those verbs. See section 22.
How do field engineers work with product and research?
Through a route with a named owner and a response commitment on the receiving side. A findings log that product is not obliged to triage stops receiving entries, and the engineers are right to stop filing. Rotation in both directions is the mechanism that keeps the route honest, because it puts people on the receiving side who remember what filing felt like.
What should be productized?
The pattern that has appeared three times, whose abstraction is stable, whose lifecycle has an owner, and whose promotion reduces cost across the portfolio rather than in the next engagement. The step organizations skip is the intermediate one: a supported extension contract where the specific thing can live without becoming core.
How do you prevent burnout?
Treat it as a structural property rather than an individual one, because the model concentrates context, customer pressure and operational load by design. Pods with real shared ownership; rotation with proven handover; on-call that includes people outside the engagement; customer load per engineer managed as a number; team-health metrics reviewed in the same meeting as delivery metrics rather than in a separate one.
What does a mature organization look like?
Marginal deployment effort falls measurably from one engagement to the next, and someone can show the series. Field findings appear in the roadmap with traceability. Partners deliver the known patterns. Access narrows over time rather than accumulating. The customization register has retirements in it. Most organizations that describe themselves this way cannot produce the first artifact on that list.
What are the main risks of adopting the model?
Over-dependence on specific individuals; insufficient governance around access and data use; and failure to route field learning into product strategy. Those three are where the serious damage concentrates, and all three are organizational rather than technical, which is usually why they are addressed last.
26 Corrections
This project began from four research documents on forward deployed engineering and enterprise ontology. Parts of them are good: the deployment-topology analysis, the challenge and mitigation pairs, the lifecycle structure and the governance discussion are all sound and are reflected throughout this page. Parts did not survive verification. Those are listed here rather than quietly dropped, because a reader has no way to tell checked work from written work unless the checking is shown.
| Claim in the source material | What verification found |
|---|---|
| “To achieve zero hallucinations, the LLM is forced to write a formal graph database query… based only on the schema.” | Not supported by any source I could find, and not claimed by the vendor most associated with the pattern. Schema conformance and correctness are measurably different: one evaluation reports 0.916 schema validity against 0.189 execution accuracy for the same model on the same tasks.[57] The phrase is quoted here only as the claim under correction; it is not asserted anywhere on this page. |
| Integrating an LLM with an ontology “shifts AI architecture from probabilistic guessing to deterministic reasoning”. | Overstated. Query execution is deterministic; question-to-query translation and result-to-prose generation are language-model inference and are not, even at temperature zero without batch-invariant serving.[74] Corrected in section 11. |
| “By integrating the LLM through an ontology, you inherently adopt a Zero Trust architecture.” | Refuted. NIST SP 800-207 defines zero trust architecture as an enterprise plan encompassing component relationships, workflow planning and access policies, and states that it “is not a single architecture but a set of guiding principles”.[21] Data-layer authorization addresses roughly three of seven tenets. Corrected in section 12. |
| “Prevented two catastrophic failures in the first year, saving an estimated $40M.” | Unattributed, and methodologically unsound regardless: prevented failures are counterfactual, so a saving derived from them is an estimate presented as a measurement. Removed. The industrial scenario in section 15 carries no outcome claim. |
| “Achieved full compliance, passed a regulatory audit with zero findings.” | No regulator, scope, period or assessor is attached. “Full compliance” is also not a state a system attains. Removed, and replaced by the vocabulary in section 22. |
| A named classified network, a three-letter agency, a twelve-facility hospital network with named clinical vendors, a North Sea platform with precise bandwidth and sensor counts. | Four unattributed narratives carrying precise, quotable numbers. Rewritten as four labelled composite scenarios with every organization, count, timeline and outcome removed. The hospital narrative also endorsed passive interception of clinical message traffic as a workaround for denied API access, which is not a practice I would recommend or repeat. |
| Security controls and audit trails “automatically generated to satisfy FedRAMP, HIPAA, GDPR, ITAR”; the edge architecture “can be pre-accredited”. | No such mechanism exists. Architecture generates evidence; organizations operate controls; assessors assess; authorizing officials accept risk. FedRAMP has also restructured — authorization is now certification, impact levels are certification classes, and the system security plan is a security decision record.[39] |
| “Secret or Top Secret/SCI clearance is mandatory, often with a polygraph.” | Nation- and contract-specific, stated as a property of the role. Not repeated. The disconnected scenario notes only that vendor remote access is unavailable in that topology. |
| A fixed team topology with an “80% code” split and universal ratios; two adoption timelines differing by a factor of two to three. | Both unattributed, and the two source documents contradict each other. This page presents phases and exit criteria without month ranges, and treats team shape as an example rather than a standard. |
| Forward Deployed Architecture described as “a coherent architectural discipline”. | I searched Open Group, IEEE, ISO, SEI, IASA, IETF and NIST material and found no recognition of the term as an architecture discipline. The certifications on offer are commercial training products without standards-body accreditation. “Forward Deployed Architect” does exist as a job title at individual companies. This page therefore avoids the abbreviation entirely and calls the design what it is: a synthesized reference architecture supporting an operating model. |
| “Kubernetes / K3s / MicroK8s: the universal application orchestration layer”; “interoperability with 100+ legacy protocols”; a policy engine that “prevents exfiltration even by privileged users”. | Three technology claims stated as requirements or guarantees. None is defensible as written. The architecture in section 04 names no mandatory technology, and the security section is careful about what a control prevents as opposed to what it detects and constrains. |
| “In late June 2026, Amazon Web Services announced a $1 billion investment” — appearing only as a truncated reference fragment. | Verified. AWS announced it on 30 June 2026 to fund a forward deployed engineering organization embedding engineers with customers for agentic AI work.[1] The claim was correct; only its context was missing. The narrower caveats are that no timeframe is disclosed for the figure, no headcount is given beyond “thousands”, and the accompanying speed claims are vendor-reported and unaudited. |
| A bibliography attributing specific titles to CNCF, Styra, Microsoft and HHS, and an institutional byline with a classification marking. | Several of those citations do not appear to correspond to real publications, and the byline and marking are not real provenance. Nothing from that bibliography is relied on here. Every source in section 27 was retrieved and read. |
Two claims I could not settle. The widely repeated figure that postings for the role grew more than 800% between January and September 2025 traces to the Financial Times through secondary coverage, but no source I found identifies the underlying dataset; it also counts postings rather than filled roles from a small base.[19] And Salesforce’s stated commitment to build a team of a thousand forward deployed engineers is a commitment, not a verified headcount.[10] The Salesforce figure appears here and nowhere else. The 800% figure appears once in section 02, carrying the same qualification.
27 Sources and method
This project combines analysis of publicly documented forward deployed operating models, architecture and security standards, peer-reviewed research on retrieval and query generation, and current AI security guidance. The architecture diagrams are reference designs rather than representations of any vendor implementation. The four deployment scenarios are composites and are labelled as such. No client, deployment, customer or measured result is described anywhere.
Method, briefly. I read the four source documents in full and built a claim-by-claim ledger classifying each substantive statement as publicly verifiable, source-derived, an architectural recommendation, a composite scenario, or an interpretation — then verified the first category against primary sources and corrected or removed what failed. Every version number, publication date, framework identifier and quoted normative sentence below was checked against the issuing body rather than written from memory, on 9 August 2026. Vendor material is labelled as vendor-reported throughout and is never presented as independent verification. Where two primary sources conflict, both are noted. Where I could not verify something, it is not asserted.
One structural caution about the subject itself. Several of the frameworks cited here changed within the last twelve months in ways that invalidate guidance still in wide circulation — among them the OWASP LLM identifiers renumbered in August 2026, a new OWASP agentic list in December 2025, SLSA v1.1 retired in November 2025, the OpenSSF baseline reissued in February 2026, FedRAMP restructured through 20x, and CMMC Phase II suspended in July 2026. Parts of this page will be wrong within a year. That is a property of the material. The fix is a re-verification cadence rather than a review date.
Forward deployed engineering: primary company sources
- Amazon Web Services, “AWS invests $1 billion to embed AI forward deployed engineers with customers”, 30 June 2026. aboutamazon.com/news/aws/aws-1-billion-forward-deployed-ai-engineers — vendor announcement. Accessed 9 Aug 2026.
- Palantir Technologies, Forward Deployed Software Engineer role descriptions, official applicant tracking system. jobs.lever.co/palantir — vendor self-description. Accessed 9 Aug 2026.
- OpenAI, Forward Deployed Engineer and Technical Deployment Lead role descriptions. openai.com/careers — vendor self-description. Accessed 9 Aug 2026.
- Anthropic, Forward Deployed Engineer, Applied AI (Federal Civilian) role description. job-boards.greenhouse.io/anthropic — vendor self-description. Accessed 9 Aug 2026.
- Accenture, “Accenture launches Microsoft Forward Deployed Engineering practice”, 18 March 2026. newsroom.accenture.com — vendor announcement; no financial figure disclosed.
- OpenAI, “OpenAI launches the Deployment Company”, 11 May 2026. openai.com/index/openai-launches-the-deployment-company — vendor announcement.
- Anthropic, enterprise AI services company announcement, 4 May 2026, subsequently branded July 2026. anthropic.com/news/enterprise-ai-services-company — vendor announcement. Note it uses “Applied AI engineers” rather than the forward deployed label.
- Databricks, AI Engineer — FDE role description, listed under Professional Services Operations. databricks.com/company/careers — vendor self-description. Accessed 9 Aug 2026.
- Vercel, Forward-Deployed Engineer role description. vercel.com/careers — vendor self-description. Accessed 9 Aug 2026.
- Salesforce, “Forward deployed engineer”, company blog, 19 November 2025. salesforce.com/blog/forward-deployed-engineer — vendor-reported; the thousand-engineer figure is a stated commitment.
Independent and analyst commentary
- Marty Cagan, “Forward Deployed Engineers”, Silicon Valley Product Group, 17 September 2025. svpg.com/forward-deployed-engineers — independent practitioner advocacy, not empirical research.
- Jina Yoon, “Forward deployed engineer”, PostHog, 11 February 2026. posthog.com/blog/forward-deployed-engineer — company blog. The on-site and infrastructure-access passage describes the industry pattern; PostHog’s own engineer is fully remote and asynchronous.
- Gergely Orosz, “Forward deployed engineers” (12 August 2025) and “The Pulse: forward deployed engineering heats up again” (24 May 2026), The Pragmatic Engineer. newsletter.pragmaticengineer.com — independent journalism. Contains no economic data; do not cite for unit economics.
- Peter Bendor-Samuel, “When are forward deployed engineers essential, and when are they not?”, Forbes / Everest Group, 30 April 2026. forbes.com — analyst-firm commentary. Contains no cost or margin data.
- Rocio Wu, “The uncomfortable truth about FDEs”, F-Prime Capital, 6 February 2026. fprimecapital.com/blog/the-uncomfortable-truth-about-fdes — venture-capital commentary. The 30–40% threshold is a stated heuristic, not a measured finding.
- Larry Dignan, “AWS launches forward deployed engineering unit”, Constellation Research, 30 June 2026. constellationr.com — analyst commentary, quoting R “Ray” Wang.
- Joe Schmidt, “Trading margin for moat”, Andreessen Horowitz, 4 June 2025. a16z.com/services-led-growth — venture-capital advocacy; a16z invests in companies using this model.
- “Forward Deployed Engineer”, Wikipedia. en.wikipedia.org/wiki/Forward_Deployed_Engineer — used only as a signpost to primary sources. Thinly sourced; several of its comparisons are editorial synthesis rather than sourced company statements.
- Financial Times reporting on growth in forward deployed engineering job postings, as relayed by PYMNTS (10 March 2026) and Fast Company (5 November 2025). The underlying dataset is not identified in any source located; the figure counts postings rather than filled roles.
- Richard MacManus, interview coverage including Sierra on the ambiguity of the term, Latent.Space, 1 July 2026. latent.space — independent technology journalism.
Standards, government and regulatory
- NIST SP 800-207, Zero Trust Architecture, August 2020. csrc.nist.gov/pubs/sp/800/207/final — final, not superseded.
- NIST SP 1800-35, Implementing a Zero Trust Architecture, NCCoE, final 10 June 2025. csrc.nist.gov/pubs/sp/1800/35/final — the earlier lettered draft volumes are withdrawn.
- NIST CSWP 29, The NIST Cybersecurity Framework (CSF) 2.0, 26 February 2024. csrc.nist.gov/pubs/cswp/29 — final.
- NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations, control catalog Release 5.2.0, 27 August 2025. csrc.nist.gov/pubs/sp/800/53/r5/upd1/final — cite the release as well as the revision.
- NIST SP 800-37 Rev. 2, Risk Management Framework for Information Systems and Organizations. nvlpubs.nist.gov — source of the assessment and authorization language quoted in section 22.
- NIST SP 800-53A Rev. 5, Assessing Security and Privacy Controls. nvlpubs.nist.gov
- NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0), 26 January 2023. nist.gov/itl/ai-risk-management-framework — final; NIST states a revision is in progress. No public draft had been issued as of 9 August 2026.
- The White House, Winning the Race: America’s AI Action Plan, July 2025. whitehouse.gov — directs the AI RMF revision; cited as the reason the framework is under revision.
- NIST AI 600-1, AI RMF: Generative Artificial Intelligence Profile, 26 July 2024. nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf — source of the “confabulation” terminology.
- NIST AI 100-2e2025, Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations, March 2025. nvlpubs.nist.gov — treats indirect prompt injection as an open problem area.
- ISO/IEC 27001:2022 including Amendment 1:2024 (climate action changes). iso.org/standard/27001 — transition from the 2013 edition closed in October 2025.
- ISO/IEC 42001:2023, Artificial intelligence — Management system; and ISO/IEC 42006:2025, requirements for certification bodies. iso.org/standard/42001 · iso.org/standard/42006
- AICPA, 2017 Trust Services Criteria for Security, Availability, Processing Integrity, Confidentiality, and Privacy (with revised points of focus — 2022). aicpa-cima.com — SOC 2 is an attestation examination, not a certification.
- Center for Internet Security, CIS Critical Security Controls v8.1, 25 June 2024. cisecurity.org/controls/v8-1
- SLSA v1.2 specification, approved 24 November 2025. slsa.dev/spec/v1.2 — v1.1 is explicitly marked retired.
- OpenSSF Open Source Project Security Baseline, release 2026-02-19. baseline.openssf.org
- HHS, HIPAA Security Rule and the proposed rule to strengthen cybersecurity of ePHI (RIN 0945-AA22), published 6 January 2025. hhs.gov · federalregister.gov · reginfo.gov — not finalized; now listed under long-term actions. The existing Security Rule at 45 CFR 164 Subpart C governs.
- PCI Security Standards Council, PCI DSS v4.0.1, June 2024. pcisecuritystandards.org — the previously future-dated v4.x requirements have been in force since 31 March 2025.
- FedRAMP 20x programme and Consolidated Rules for 2026. fedramp.gov/20x — Phase 3 active from April 2026; terminology and artifacts renamed.
- US Department of Defense CIO, CMMC programme status including the suspension of Phase II requirements announced 13 July 2026; DFARS Case 2019-D041 final rule effective 10 November 2025. dodcio.defense.gov/CMMC · federalregister.gov
- NIST SP 800-171 Rev. 3, Protecting Controlled Unclassified Information in Nonfederal Systems and Organizations, 14 May 2024. csrc.nist.gov/pubs/sp/800/171/r3/final — note that CMMC Level 2 is still assessed against Rev. 2.
AI and application security guidance
- OWASP, LLM01 Prompt Injection, OWASP Top 10 for LLM Applications. owasp.org — source of the statement that fool-proof prevention is unclear.
- OWASP GenAI Security Project, OWASP Top 10 for LLM Applications 2026, August 2026. genai.owasp.org — identifiers renumbered from the 2025 edition; Excessive Agency is now third. OWASP’s own legacy page still displays 2025 identifiers.
- OWASP GenAI Security Project, OWASP Top 10 for Agentic Applications, December 2025 (ASI01–ASI10). genai.owasp.org
- OWASP MCP Top 10. owasp.org/www-project-mcp-top-10 — Incubator project at v0.1 beta; cite with less weight than a flagship document.
- MITRE ATLAS, data release v2026.05, May 2026. atlas.mitre.org · github.com/mitre-atlas — continuously updated; techniques now tagged by platform including agentic AI.
- Dane Stuckey, Chief Information Security Officer, OpenAI, public statement on agent-mode security, 21 October 2025. Vendor statement — “prompt injection remains a frontier, unsolved security problem”.
- Anthropic, “Prompt injection defenses”, 24 November 2025. anthropic.com/research/prompt-injection-defenses — vendor-reported; states a 1% attack success rate “still represents meaningful risk”.
- Greshake, Abdelnabi, Mishra, Endres, Holz & Fritz, “Not what you’ve signed up for: compromising real-world LLM-integrated applications with indirect prompt injection”, ACM AISec 2023, arXiv:2302.12173. The canonical citation for the indirect prompt injection class.
- Model Context Protocol specification, revision 2026-07-28, including Security Best Practices. modelcontextprotocol.io/specification — source of the quoted normative requirements.
- CVE-2025-49596, MCP Inspector unauthenticated remote code execution, CVSS v4.0 base score 9.4, published 13 June 2025. nvd.nist.gov/vuln/detail/CVE-2025-49596
Research
- Palantir, Ontology documentation and core concepts. palantir.com/docs/foundry/ontology — vendor documentation; no independent technical evaluation located.
- Palantir, “Reducing hallucinations with the Ontology in Palantir AIP”, 8 July 2024. Vendor engineering blog. Note the verb: reducing, paired with human oversight.
- Edge, Trinh, Cheng, Bradley, Chao, Mody, Truitt, Metropolitansky, Ness & Larson, “From local to global: a Graph RAG approach to query-focused summarization”, arXiv:2404.16130. Microsoft Research. Headline result concerns comprehensiveness and diversity, evaluated by preference judgement, not factual accuracy.
- Microsoft, GraphRAG responsible AI transparency documentation. github.com/microsoft/graphrag — states that expert human verification of answers is needed and that the system is designed for trusted users.
- Chauhan, Raj, Mujumdar, Saha & Jain, “Mind the query”, EMNLP 2025 Industry Track. aclanthology.org — peer-reviewed; 76.75% execution accuracy overall for the strongest model tested, 49.93% on complex aggregation.
- Ranganath & Raghavendra, “PIPE-Cypher”, arXiv:2606.08481, June 2026. Preprint, not peer-reviewed. Reports 0.916 schema validity against 0.189 exact execution accuracy.
- Shen, Wan et al., “Understanding, detecting and repairing real-world in-context-learning-based text-to-SQL errors”, PACM SE (FSE), arXiv:2501.09310. Peer-reviewed; semantic errors account for 36.1% of errors on one benchmark.
- BIRD text-to-SQL benchmark leaderboard. bird-bench.github.io — accessed 9 Aug 2026; human baseline 92.96% execution accuracy.
- Niu et al., “RAGTruth: a hallucination corpus for developing trustworthy retrieval-augmented language models”, ACL 2024. aclanthology.org — peer-reviewed.
- Jacovi et al., “The FACTS Grounding leaderboard”, arXiv:2501.03200. Google DeepMind — vendor-reported research.
- Joren, Zhang, Ferng, Juan, Taly & Rashtchian, “Sufficient context: a new lens on retrieval augmented generation systems”, ICLR 2025, arXiv:2411.06037. Peer-reviewed.
- Xu, Jain & Kankanhalli, “Hallucination is inevitable: an innate limitation of large language models”, arXiv:2401.11817. Preprint, widely cited and contested.
- Kalai, Nachum, Vempala & Zhang, “Why language models hallucinate”, arXiv:2509.04664; OpenAI, 5 September 2025. Vendor-reported. Argues confident hallucination is reducible through abstention while accuracy never reaches 100%.
- Huang, Chen, Mishra, Zheng, Yu, Song & Zhou, “Large language models cannot self-correct reasoning yet”, ICLR 2024, arXiv:2310.01798. Google DeepMind; peer-reviewed.
- Peeters, Bizer et al., “WDC Products: a multi-dimensional entity matching benchmark”, EDBT 2024, arXiv:2301.09521. Peer-reviewed; 89.04 F1 on seen entities against 64.56 on unseen.
- Agrawal, Kumarage, Alghamdi & Liu, “Can knowledge graphs reduce hallucinations in LLMs? A survey”, NAACL 2024, arXiv:2311.07914. Peer-reviewed; frames the contribution as mitigation.
- Ozsoy, “Enhancing Text2Cypher with schema filtering”, arXiv:2505.05118. Neo4j — vendor-reported research; benefit was not uniform across models.
- METR, “Measuring the impact of early-2025 AI on experienced open-source developer productivity”, 10 July 2025, arXiv:2507.09089. Randomised controlled trial, 16 developers, 246 issues. A snapshot of early-2025 tooling on a specific kind of work.
- Zhu, Jin, Pruksachatkun et al., “Establishing best practices for building rigorous agentic benchmarks”, arXiv:2507.02825. Preprint.
- Liang, Bommasani, Lee et al., “Holistic evaluation of language models (HELM)”, TMLR 2023, arXiv:2211.09110. Peer-reviewed; seven evaluation axes.
- Ong, Almahairi, Wu, Chiang, Wu, Gonzalez, Kadous & Stoica, “RouteLLM: learning to route LLMs with preference data”, ICLR 2025, arXiv:2406.18665. Peer-reviewed; reported savings vary by an order of magnitude across benchmarks.
- Yan, Bischof, Frye, Husain, Liu & Shankar, “What we learned from a year of building with LLMs”, O’Reilly Radar, 28 May 2024. Practitioner synthesis, vendor-neutral.
- He, “Defeating nondeterminism in LLM inference”, Thinking Machines Lab, 10 September 2025. Vendor-reported but reproducible; 1,000 completions at temperature zero produced 80 unique outputs before batch-invariant kernels were applied.
- Turan, “Oversight has a capacity: calibrating agent guards to a subjective, fatiguing human”, arXiv:2606.08919, June 2026. Single-author preprint, not peer-reviewed. Cited for direction, not for its specific figures.
- Parasuraman & Riley, “Humans and automation: use, misuse, disuse, abuse”, Human Factors 39(2), 1997. The canonical source for automation misuse and complacency; cited for the concept.
Discuss enterprise architecture. If your organization is standing up a forward deployed capability, or working out whether an AI deployment is ready to be called production, I am happy to talk through how this reference translates to your environment. Get in touch or explore more projects.