Skip to main content
AI Adoption11 min read·Verified Sep 28, 2026 against first-party sources

Should IT Build Every AI Agent, or Train Domain Experts to Build Their Own?

The central build team does not stall on capacity. It stalls on specification. Autonomous domain experts do not stall on prompting skill. They stall on the things a platform never reports. The agent is not the unit of ownership. The layer is.

MK
Mathieu Kessler
Founder, Kesslernity

The pilot worked. Somebody built an agent that saves one team six hours a week, and now forty other teams want one. So the question lands on somebody's desk: does IT build all of them, or do we teach the business to build their own?

Both answers are wrong, and they are wrong in different places. That matters, because the two failure modes are not symmetrical, and the fix is not a compromise between them.

Why the central build team stalls

The obvious objection to “IT builds everything” is capacity. That objection is out of date.

Building an agent is no longer the expensive part. A declarative agent is a name, a description, an instruction block and a set of knowledge sources. Microsoft caps the instructions at 8,000 characters, and caps the knowledge per source type: 100 SharePoint files, folders or sites, 50 OneDrive files, four public website URLs, five Teams chat URLs. Limits shaped like that tell you how much of this is configuration rather than engineering. An afternoon, most of the time.

A prototype, that is. Running a dependable service is a different job, and the distance between those two is most of what this post is about.

What a central team cannot supply is the definition of a correct answer.

Ask a platform engineer to build an agent that reviews a method statement, or reconciles a schedule against a contract, or screens a vendor document package for missing deliverables. They can build it. They cannot tell you whether the output is right. Neither can the model. Only the person who has done that job for eleven years can read an output and say “this is confidently wrong in paragraph three, and here is the standard it is quoting from an edition that was withdrawn.”

So the central team becomes a queue. The queue is where these programmes die, and slowness is the least of it. Every item in that queue needs a domain expert's time anyway, on the specification side. You have added a ticket, a handoff and a two-week round trip to a conversation two people could have had directly.

You can spot this failure mode without looking at the backlog. Look instead for agents that work and nobody uses. They pass IT's tests and fail the professional's judgement, and nobody can say precisely why, because the person who could say why was last in the room at kickoff.

None of which argues that a central team should stop writing software. It should. Somebody has to own retrieval quality, the regression tests that run when the model changes, integration testing, and the pager at two in the morning. None of that is a domain expert's job and none of it is optional. The distinction worth holding is between engineering the service and specifying the work. A central team building against acceptance criteria the experts wrote is doing ordinary software delivery, and it works. A central team also handed the question of what a correct answer looks like is the one that turns into a queue.

Why autonomous domain experts stall

So try the other answer. Train the experts, let them build, get out of the way.

This one fails somewhere most adoption programmes never look. Domain experts cannot see what the runtime does when nobody is watching, and no amount of prompt training will teach them. Four examples, none of them about prompting.

The guardrail that is quietly off. Mistral's Vibe CLI lets you attach hooks that gate what the agent is allowed to do. In version 2.25.5, released 18 September 2026, those hooks fail open by default: if the hook itself errors, the action proceeds. You have to set strict to make a broken guardrail block. A domain expert who writes a hook has every reason to believe the guardrail is on. It is on right up until the day it breaks, and the day it breaks is the day it stops telling you anything.

The grounding limit nobody reads. Microsoft publishes hard ceilings on what SharePoint actually indexes for Copilot. Above 150 MB for most file types, or 512 MB for PDF, PPTX, PPT, DOC and DOCX, SharePoint downloads only the document metadata, not the full content. Parsing stops at 2,000,000 characters, and SharePoint marks the item partially processed. The word breaker stops at 1,000,000. There is a 30-second budget for parsing a single item and its attachments, and a second 30-second budget for word breaking. My read is that depth is partly a function of parse time, not only of what is in the file. The practical version: an agent built over a document library looks perfect in testing on the small files and silently thins out on the big ones. An item shared with more than 10,000 distinct users or Entra security groups stops being searchable by any user. It survives only as an eDiscovery result.

None of that shows up as an error. It shows up as an answer that is merely incomplete, which is the hardest failure in this whole field to catch.

The bill. The Copilot seat does not cover what the agent does. Custom agents that ground in Microsoft 365 data through the Work IQ APIs are billed in Copilot Credits at $0.01 a credit on pay-as-you-go, and Microsoft is explicit that this use is not an entitlement of the Copilot licence. Capacity packs run $200 per pack per month for 25,000 credits, billed annually, and unused credits do not roll over. Agent 365, at $15 per user per month on an annual commitment, buys the identity and governance layer on top of that; Microsoft's own wording is that there are no consumption-based costs for Agent 365 yet. The rest arrives from elsewhere, including Windows 365 for Agents as a separate purchase if your agents need a machine. So an enthusiastic expert with a build button can create recurring spend that lands on someone else's cost centre, and the person who could have set a ceiling was never told the agent existed.

The floor moves. Mistral shipped 58 releases of Vibe between 2.5.0 in March 2026 and 2.25.8 in September 2026, and three of those landed inside a single week in September. Read the 2.25.5 source and you find one of its rollout flags cached with a seven-day time to live, which means the same machine, running the same version, with nobody touching anything, can behave differently next week. Ask a process engineer to own that and you are asking them to do a job nobody described to them when they volunteered.

Call this what it is: platform opacity. It belongs to the platform team to close, because nobody else can see it.

The split that actually holds

Stop asking who builds the agent. The agent is not the unit of ownership. The layer is.

IT owns the rails, not the agents. Identity, data boundaries, approval gates, spend ceilings, logging, lifecycle, and a kill switch that somebody has actually tested. Add the one most estates forget: what happens to an agent when its named owner changes role or leaves. A review date is a note in a calendar, and calendars do not enforce anything. A forcing function is the agent going read-only, then off, when the owner field stops resolving to an employee. Build the second one. The first one has never once fired on its own.

Domain experts own the specification and the verdict. What the agent is for, what a correct output looks like, and whether this particular output is correct. None of that can be delegated upward, and none of it can be trained into IT, because it is not knowledge. It is judgement built from doing the work.

That comes with a caveat, and it is the one that breaks most versions of this model. A verdict only catches what a reader can see. The grounding limits above do not produce wrong answers, they produce quietly thin ones, and an expert reading carefully will pass them. So the verdict is necessary and it is not sufficient. What closes the gap is sampling: a set of cases with known right answers, re-run whenever the model, the connector or the content changes, with the results read by a human. The expert writes the cases, because only they know the right answers. The platform team runs them, because only they know when something changed. That is the single obligation in this whole split that belongs to both groups, and it is reliably the first thing cut.

Nobody is autonomous. Autonomy is the wrong word for what a domain expert actually needs: a bounded envelope they can move quickly inside without asking permission for every step. Autonomy without rails gives you four hundred orphaned agents, half of them pointing at data nobody audited. Rails without a fast lane give you the queue. You need both, and they are different people's jobs.

The uncomfortable implication for most org charts: this makes your best domain experts part-time product owners, and it turns your platform team into a service rather than a build shop. Both cost money. I am not going to hand you a ratio of platform engineers to agents, because anybody quoting one has not run it. Three things set that number: how many agents sit in the fast lane, how often the platform changes underneath them, and whether the sampling above is a funded job or an aspiration. Size those and you have your answer. Skip them and you have dismantled the build-shop budget without replacing it, which is worse than either answer to the original question.

Super agents are the wrong unit

One more thing, because “super agent” is doing a lot of work in vendor decks this year.

A single agent that runs a whole workflow end to end is the least auditable object you can build. When a chain hands off internally, a step that is missing an input will often proceed on a plausible substitute rather than stop. The final output still looks finished. It is formatted, confident and complete, and the defect is four steps upstream where nobody will look, because nothing failed.

Smaller agents, with a named human owner each and a checkpoint wherever the next step assumes the previous one was right, are slower to demo and dramatically cheaper to trust. If you cannot say which human reads the output between step two and step three, you have not built an agent. You have built a way to launder an assumption into a deliverable.

Decision tree 1: who builds this one?

Run any request through five questions, in order. Stop at the first one that changes the answer.

  1. Does anything downstream authorise work, or sign off on safety? Permit to work, lockout tagout, confined space entry, a job safety analysis, an incident classification, an inspection sign-off. Be precise here, because the usual slogan can be read two incompatible ways and a policy has to pick one. An agent may assemble, draft and check the material that feeds one of those decisions. It may never issue the authorisation, approve it, or stand as the record of it. AI prepares, humans decide, and the signature is the line.
  2. Does it write, or only read? Read-only is the fast lane, subject to the next question. Anything that writes, sends, books, files or posts needs an approval gate and a rollback story before it needs a prompt.
  3. Does it cross a data boundary, and is that boundary real? A new connector, site or system of record is the easy case: rails decision, IT's call, not negotiable by enthusiasm. The trap is the other case. “Data the person already has access to” sounds like a safe boundary and is not, because the standing failure in this product class is surfacing content from overshared locations a user can technically reach and was never meant to find. An agent searches that surface far more thoroughly and far more patiently than the person ever did. Somebody runs the oversharing report before the fast lane opens, rather than assuming it.
  4. One person, or a process? One person's speed is theirs to build, today, no ticket. The moment a second team depends on the output, it needs a named owner, a version and a review date, and then it is a joint build.
  5. Will somebody have to explain this in eighteen months? If yes, it cannot live in a chat window. It needs to be a versioned artifact in source control with the prompt, the knowledge sources and the change history attached.

Anything that clears all five is the expert's to build this week. That is usually more requests than a platform team expects, which is the point.

Decision tree 2: which runtime, and does it matter?

It matters less than the vendor comparison tables suggest, and the reason is worth internalising: the job travels, the configuration does not.

  1. Where does the data already live? Pick the runtime that already has a governed path to it. Moving data to reach a preferred model is how a tool choice turns into a compliance project.
  2. Is the output a document, a decision or an action? Documents are the easy case on every platform. Decisions need the source visible next to the answer. Actions need the gate from tree 1, question 2, and the gate has to live in the platform, not in the instructions.
  3. Who maintains it when the model changes? Not if. Every runtime here has replaced its default model at least once this year. If the answer is "the person who built it," you have a single point of failure with a notice period.
  4. Is it vendor-specific work, or portable work? Governance, review gates and the definition of a correct answer are portable and will survive the next procurement decision. Connector configuration, instruction syntax and CLI internals are not, and they expire faster than most people budget for.

Inside Microsoft 365 specifically, question 1 has a sharper version, because Agent Builder, Copilot Studio and Foundry have different data paths and different bills. That one has its own tree, free on this site: Which Microsoft AI agent tool do you actually need?

If you are the one who has to run this

The rails do not build themselves

Everything above is the operating model. These are the artifacts that implement it: the rollout plan, the governance gates, the agent templates, and the cost model for the metered spend this post describes. Four doors, depending on where you are.

The bottom line

Nobody actually needed to settle whether IT should build the agents. The question underneath it is which layer each group owns, and that answer holds across every runtime you will run this decade.

IT owns the rails. Domain experts own the specification and the verdict. Both own the sampling that catches what neither would see alone. Nobody is autonomous, and nothing goes near a safety authorisation.

If you get that split right, the tooling decision is a detail, and it stops being an argument.

Verified on 28 September 2026. SharePoint indexing limits from SharePoint search limits (updated 25 June 2026). Declarative agent limits from Agent Builder knowledge sources and the declarative agent best practices guide. Pricing from the Copilot Studio licensing guide, the Copilot Credits guide, and the Agent 365 licensing FAQ. Vibe behaviour was read against version 2.25.5, released 18 September 2026; the current release at the time of writing is 2.25.8, published 23 September 2026, so check the version you are actually running. Prices are US list and change. Kesslernity is not affiliated with, or endorsed by, Mistral AI, or by Microsoft.