A working paper · DOC 001

Discovery is the bottleneck

Why the scarce skill in applied AI is no longer answering questions, but knowing which ones to ask. A thesis on auditable judgment and the engine that compounds it.

Brendan Sibeth, founder and managing partner · v1.0 · 2026

Abstract

The market is racing to automate answers. We argue the binding constraint sits one step earlier, in deciding what is worth asking, and that the durable advantage in applied AI accrues to whoever can do three things in sequence: find the questions worth asking, encode how experts actually reason, and prove every resulting decision back to its source. We describe that system as a loop rather than a pipeline, show why it only compounds once it observes its own outcomes, and explain why institutional finance is where such a system is worth the most and, counter-intuitively, the last place you should try to teach it.

01

The inversion

Almost every tool in the current wave begins at the same place: it assumes the question is given and races to produce the answer. That is a reasonable bet, because answering has become cheap. Large models have commoditized the step the last decade treated as hard. But the commoditization of answering exposes what was always the deeper problem, and it is not a problem of answers at all.

The real bottleneck is that, inside most valuable domains, nobody knows which questions to ask. A problem that is never named never gets the conversation that would solve it. Expertise sits on top of fragmented, unstructured knowledge that the people closest to it have stopped noticing, because they have accepted the chaos as normal. The industry solves given a question, find the answer. Almost no one solves given a pile of messy domain knowledge, what are the right questions in the first place?

Get the questions right and the knowledge graphs, the reasoning chains, and the decision architecture all become buildable on top of them.

This is not a rhetorical flourish. It is the load-bearing claim of this paper, because it is also the most leveraged one. Get the questions wrong and no amount of model quality rescues the result. The advantage belongs to whoever owns the step before the answer.

02

Why now, and why this failed before

A careful reader will object that this idea is old. They are right. Externalize expert reasoning into a reusable, machine-readable form is the expert-systems dream of the 1980s, the knowledge-management programs of the 1990s, and the ontology engineering that followed. Each died on the same two rocks. First, tacit knowledge resists being told: experts know more than they can articulate, so asking them to write down how they decide produces a thin, lossy caricature. Second, elicitation did not scale. The field literally named this the knowledge-acquisition bottleneck, which is our what to ask problem wearing a lab coat thirty years early.

So the honest question is not whether the idea is good. It has been good, and fatal, for decades. The question is what has actually changed. The answer is precise: the two failure modes that killed every prior attempt are exactly the two that recent capability shifts have dissolved. Elicitation can now be conversational rather than a consultant with a whiteboard, and a model can interview an expert and follow the hesitation.

03

A loop, not a pipeline

The system has four moves over one shared spine.

Discovery

Finds which problems are worth solving.

Elicitation

Finds which questions actually define a problem once you are inside it.

Encoding

Captures how an expert reasons, not merely what they conclude, as portable primitives a new agent can inherit.

Provenance

Gives every decision its DNA: the source it drew on, the reasoning it followed, the confidence it carried, the decision it reached, and the outcome that resulted.

Most platforms bolt explainability on at the end. Here it is the foundation, because in regulated domains a decision you cannot defend is a decision you cannot make.

Drawn as a loop, a property appears that a pipeline never has: the provenance layer's own low-confidence decisions are a live map of where the right questions have not yet been asked. Uncertainty, recorded honestly, becomes the discovery engine's next input. The output of the back end is the fuel of the front end. That circulation is the difference between a system that runs and a system that compounds.

04

The flywheel needs a teacher

There is a subtle failure waiting inside that loop, and naming it is the most useful thing this paper can do. A loop that records confidence is not yet a loop that learns. If the only signal that circulates is the model's own confidence, the system can grow more and more certain of reasoning that is wrong, and nothing ever corrects it. That is an echo chamber with excellent provenance.

A loop that records confidence is not yet a loop that learns. Confidence is a self-assessment; correctness requires outcomes.

Closing the loop honestly requires a teacher: the ground-truth outcome of each decision and, cheaper and sharper, the moments a human in the loop overrides the recommendation. An override marks precisely where the encoded reasoning and a trusted human diverged, which makes it among the most valuable signals a judgment system can capture. Harvest those, and the trust boundary stops being a place where feedback leaks out and becomes the place where it comes in.

This has a hard architectural consequence. To learn, the system must sit where decisions and their consequences both pass through it. It must be the system of record, not an adviser on the side. An advisory layer that hands over a recommendation and never sees what happened is structurally cut off from its own report card. It can be useful; it cannot improve.

05

Learn where it is cheap; deploy where it is dear

The domain in which auditable judgment is worth the most, institutional finance, is the domain in which the learning loop turns slowest. A compliance or allocation decision is not adjudicated by reality for quarters, sometimes only at an audit, and sometimes never cleanly. When returns do arrive they are so confounded by market noise that separating a good judgment from a lucky one is its own hard problem. You cannot debug a learning loop on a multi-quarter, noise-soaked delay.

The resolution is to separate where you learn from where you earn. Calibrate the engine in a domain where ground truth returns in days, an operational setting that confirms or refutes a call almost immediately, and carry the proven machinery into finance, where each decision is worth far more and must be provable. In finance the value was never autonomy in the first place. It is a provenance spine that defends every decision to a board or a regulator, and a policy layer that re-runs the same business day a rule changes.

06

What an institution is actually buying

Everyone will have automation; it confers no advantage to have what everyone has. The durable assets are different. The first is a method for externalizing the judgment that today lives tacitly in your most experienced people, judgment that walks out of the building when they do. The second is a record that proves every decision back to its source, on demand, to whoever is entitled to ask. The third is a discovery engine that turns your own institution's uncertainty into its next set of questions, so the system gets sharper precisely where it is currently weakest.

The moat, in other words, is not the model. It is the encoded reasoning and the provenance that compounds across decisions: a portable protocol for how a domain thinks, defended by a spine that makes every output auditable. That is a far harder thing to copy than a prompt, and a far more valuable thing to own than another layer of automation.

On the limits of this thesis

Stated plainly, because a thesis that hides its soft points is not worth defending.

The compounding is bounded, not magical.

It rests on a repeatable method and on structure demonstrated across real systems, not on a claim that all domains converge. We reject the version that would promise more than the evidence supports.

Elicitation is the live frontier.

Our bet is that recent capability shifts changed its economics, not that the problem is solved. The knowledge-acquisition bottleneck can reappear in new clothes, and we build against that.

Outcome latency is real.

In slow-feedback domains the loop is calibrated elsewhere and carried in. Where we cannot yet observe ground truth quickly, we say so, and we do not dress confidence up as calibration.

The argument is above in full, open to read and quote. The typeset PDF, with the exhibits and the notes, is delivered on request.

Part of the document corpus · Sheet 07