The harness is the product. Here is its input/output spec.
Nine in ten agent pilots never reach production, and the engineering is rarely what fails — the inputs are missing. Every brief, journey map, and measurement plan a strategist has ever written is a named component of the agent harness. Each one comes back out as a number somebody can audit.
An agent harness is everything in an agent that is not the model — prompt construction, tools, memory and state, permission boundaries, validators, and the loop that decides whether any of it worked. It is also where every deliverable a strategist has ever produced goes to become executable.

Part 1 argued that the scarce half of forward deployment is the strategist seat: reading the business, defining what the system is for, owning whether it gets adopted. This is the follow-through. What does that seat actually hand over, and what comes back?
The crossing almost nobody makes
Deloitte's Tech Trends 2026 puts 38% of organizations piloting agents and 11% running one in production. Gartner still expects more than 40% of agentic projects to be cancelled by the end of 2027. The interesting part is what leaders blame.
Evaluation gaps (64%), governance friction (57%), model reliability (51%). Not one of those is a model-capability problem. All three are harness problems, and all three trace back to an input nobody wrote down before the build began. Meanwhile the agents that do cross return an average 171% ROI, with a median payback under nine months. The gap is not technical difficulty. It is unspecified work.
The model is the cheapest part of the system. The expensive part is deciding what it is allowed to do, and what would count as it working.
The input spec
Fifteen years of brand, digital, CX, and data strategy produced four artefacts on repeat: a brief, a journey map, a measurement plan, and a segmentation. In an agentic system, none of those are background documents. Each one is a component with a slot.
The brief is the system prompt and the permission boundary — what this thing is for, and what it may never say or do. The journey map is the tool allowlist and the escalation gate — where the agent acts unattended, and where a human takes the wheel. The measurement plan is the eval harness. The segmentation is the retrieval strategy.
Journey mapping was always a measurement discipline wearing a design costume. You were not documenting where customers went. You were documenting where value was created, where it leaked, and which single decision at which single moment would have changed the outcome. Those are precisely the decisions an agent has to encode.
The output spec
Here is the part that makes this concrete rather than tidy: the outputs are numbers, and they are now benchmarkable. Microsoft's AgenticRAG paper (May 2026) gave a model search, open, and summarize tools instead of one-shot retrieval, and measured what came out — 49.6% recall@1 on BRIGHT, 21.8 points over the best embedding baseline, and 92% answer correctness on FinanceBench, within two points of having oracle access to the true evidence.
Read that as a strategy result, not an engineering one. The gain came from changing how the system is allowed to look for evidence — a retrieval strategy decision, made upstream, by whoever defined what "relevant" means for this business. The model was held constant. The harness moved the number.
That is the whole argument in one benchmark. Retrieval quality, faithfulness, tool-selection accuracy, escalation rate: these are the measurement plan, executed. A strategist who could defend a media-mix model to a CFO can defend these, because they are the same object.
Where the harness carries legal weight
In regulated deployments each tier picks up an obligation. Runtime needs audit trails and encryption. Capabilities need the tool list and the retrievable corpus governed against a real framework — NIST AI RMF, or FDA and MLR review in pharma. Assurance needs policy-as-code and human approval before anything irreversible.
We build this way because we have had to. For a Fortune 500 manufacturer we encoded a lead scientist's relevance rules — what counts as a real commercial signal, and what is noise — before a single pipeline stage was written; the rules were the product, and the pipeline was how they ran at volume. For a pharma AI Center of Excellence, the claims and governance layer was written first, and the agents were built inside it rather than retrofitted to it.
In both, the sequence was the same and it is the only part worth copying: the inputs were specified, then the system was built to satisfy them. Teams that reverse the order end up in the 89%.
What to do about it
This is Part 2 of The Forward Deployed Strategist, a four-part series. Next: the economics — who staffs this seat, what it costs, and why the pricing model that built the agency business does not survive contact with agentic delivery.
Powered by Enso Labs
Frequently Asked Questions
What is an agent harness?
Everything in an agent that is not the model — prompt construction, tools, memory and state, permission boundaries, validators, and the evaluation loop. It is the software layer that turns a stateless model into a system that can run a task end to end. Practitioners group it into three tiers: Runtime, Capabilities, and Assurance.
Why do most AI agent pilots fail to reach production?
Not because the model is weak. Deloitte's Tech Trends 2026 finds 38% of organizations piloting agents and only 11% running one in production. Leaders name evaluation gaps (64%), governance friction (57%), and model reliability (51%) as the blockers — all three are harness problems, and all three trace back to inputs nobody defined before the build started.
How does a strategy deliverable become part of an agent?
Directly and one-to-one. The brief becomes the system prompt and the permission boundary. The journey map becomes the tool allowlist and the escalation gate. The measurement plan becomes the eval harness and its golden-set CI gates. The segmentation becomes the retrieval strategy and its recall target. Same craft, executable form.
What does Enso Labs do before building an agent?
Enso Labs writes the inputs first — the outcome the agent serves, the actions it may take unattended, the point it must escalate to a human, and the metric that will prove it worked. Only then does the build start. Get in touch at https://ensolabs.ai/contact.
Want to scope an engagement around this?