The Model is Only Half the Story. Let’s Talk about Your Institutional Knowledge

The Model is Only Half the Story. Let’s Talk about Your Institutional Knowledge

Chandini Jain

Industry Insights

If an AI workflow produces polished, plausible work with the wrong institutional standards, is it still useful? 

If a new employee did the same, the answer would probably be “no,”  And more experienced colleagues would train them, review the work, and explain the exceptions and nuances of why certain decisions were made. 

As models become better at building and maintaining agents and workflows, business leaders may assume much of that institutional effort is irrelevant. The engineering burden may fall dramatically, but the management burden does not.

Anthropic's own research on agent autonomy makes the distinction clear. Software engineering is comparatively amenable to autonomous work because code can be tested before release. In finance, law, and other domains where verifying an output may require the same expertise as producing it, the transition may be slower or take a different form. Anthropic, February 2026

This reveals a critical constraint that all leader should be assessing as they develop their AI strategies.

The model is relatively useless without institutional knoweldge 

Models will increasingly write integrations, construct workflows, generate tests, diagnose failures and implement repairs. Model providers, cloud companies, and enterprise platforms will supply more of the routing, monitoring, and orchestration beneath them.

But building the workflow is not the same as knowing what the workflow should do.

In our deployments, some of the hardest work happens before a reliable evaluation can even be written. Our operating experts have to work deeply with a firm's teams to extract the knowledge that lives in people’s heads. Work like knowing which sources they trust, how they treat contradictory evidence, what makes an exception material, which edge cases require escalation, and what a senior reviewer recognizes as excellent work.

Much of this knowledge is not in the firm's data lake or methodology documents. It is tacit, contested, and idiosyncratic. Different experts may follow different rules without realizing it until they are forced to agree on what the agent should do.

No foundation model begins with access to that knowledge. A model can help elicit it. It can propose a rubric, generate test cases, and grade routine outputs. But it cannot be the sole authority on whether its own work reflects the institution's standard—particularly when the institution itself has never made that standard explicit.

Anthropic's guidance on agent evaluations reflects this boundary: model-based graders need calibration against human experts, while complex domains such as finance require expert judgment for subjective or ambiguous work. Anthropic, January 2026

Without that institutional specification, an agent can be technically functional and operationally incompetent. It may find the right documents and calculate the right ratios, yet apply the wrong materiality threshold. 

Putting that agent into production is the equivalent of hiring an incompetent employee, giving them access to institutional systems and allowing their work to scale instantly.

A specialist partner must do more than build

This changes what a financial institution should demand from an AI partner.

The partner's value cannot rest primarily on having better access to models or proprietary agent infrastructure. Both will commoditise. Nor is it enough to deliver a workflow and return responsibility to the client.

A credible specialist should perform three continuing roles.

1. Enable the institution to define what “good” looks like.

A small group of senior business owners and subject-matter experts must ultimately define and approve what acceptable work means. A strong partner will not allow this responsibility to remain diffuse. 

2. Turn that standard into an operating playbook.

The partner must translate examples, exceptions and institutional judgment into workflow logic, evaluations, escalation rules and evidence requirements. This is how knowledge held by a few experienced people becomes reusable operating capacity.

3. Remain accountable for performance over time.

The work does not end at first deployment. Models change, policies evolve and new edge cases appear. A specialist partner should operate the workflow, judge outputs against the institution's approved standards, investigate failures and keep the system reliable at a predictable cost.

Models will improve, but the context still matters 

Perhaps the model will soon do most of this too.

It will interview subject-matter experts, infer rules from historical decisions, identify inconsistencies, generate evaluations and monitor its own outputs. Firms with strong internal leaders may be able to use those capabilities without a specialist partner.

But better automation does not eliminate the need to decide whose judgment counts or to reconcile disagreements. Those are institutional decisions that don’t just disappear with a better model.

As engineering becomes abundant, the scarce capability will be turning institutional judgment into an explicit, testable, and continuously operated standard of work. Firms that do this well will accumulate institutional memory and expand expert capacity at an unprecedented pace. 

And those who choose to keep the playbooks locked in the heads of senior leaders? 

It won’t be a matter of trying to keep pace or catch up. Their competitors will be playing in a completely different game.

See what agentic AI does for your team

A 15-minute demo focused on your workflows, not a generic product tour.