An LLM Is a Judgment Engine, Not a Function Runner

If a deterministic algorithm can do the job, an LLM shouldn't — it will be slower, costlier and less reliable. Here's the test we apply, and where language models genuinely earn their place in market work.

Published August 17, 2026

There is a simple test we apply before letting a language model anywhere near a piece of our pipeline:

If you can write the assertion, write the function.

If you can state, in advance and without ambiguity, what a correct output looks like — the 10-year yield is 4.28%, the 50-day moving average crossed the 200-day, this quarter's revenue grew 6.2% — then a deterministic algorithm can produce it, and should. Reaching for an LLM at that point is not innovation. It is choosing a slower, more expensive, less reliable version of something you already know how to compute.

That sounds obvious written down. It is also the single most common way we see AI wasted in this industry.

The case against the LLM as function runner

When you use a language model to do something a function could do, you inherit every one of its weaknesses and none of its strengths.

It is non-deterministic by construction. The same input can produce different output. For a market brief, that is not a rounding error — it is a different number in front of someone about to size a position. A function that computes a yield spread returns the same answer every time, forever.

It fails silently and plausibly. A broken function throws. It gives you a stack trace, a line number, a test that goes red. A language model asked to compute something it cannot compute does not throw — it produces a confident, well-formatted, wrong number that looks exactly like a right one. There is no exception to catch. The failure mode of arithmetic is a crash; the failure mode of a language model is a fluent lie.

You cannot unit-test it the way you test code. You can evaluate it statistically, across a sample, with a tolerance. That is a genuinely useful discipline and we do it. But it is not the same as asserting that spread(4.28, 3.91) == 0.37 and knowing it will hold on every run until someone edits the function.

It is orders of magnitude more expensive and slower. A moving average is a handful of floating-point operations. Asking a model to "calculate the 50-day moving average of these closes" is a network round trip, a few thousand tokens, and a wait — to get a worse answer.

It cannot be audited. When a regulator, a subscriber, or your own future self asks why does it say 4.28%, the honest answer for a function is a line of code and a data source. For a model, the honest answer is a shrug.

None of this is an argument against language models. It is an argument about where they belong.

What LLMs are actually extraordinary at

A language model is a judgment engine. It is worth reaching for exactly when the task has no single correct answer that you could have specified in advance — when the output depends on taste, salience, framing, or synthesis across things that were never in the same table.

Put plainly: use it where you would otherwise need a person.

Here is where that lands in market work.

1. Deciding what the story is

On any given morning there are two hundred things that happened overnight and room for maybe six. Which six? That is not a ranking problem you can solve with a sort key. It depends on what has already been covered, what a reader is carrying from last week, what has changed versus what has merely continued, and what a specific audience already knows.

We run an explicit editorial pass before anything is written, and its instructions begin by saying what it is not: this pass is not writing, it is deciding what the ongoing story is. That is a judgment call, made fresh every day, against a context no schema anticipated. A language model is the right tool. A SELECT ... ORDER BY is not.

2. Turning verified numbers into a sentence a human wants to read

This is the highest-value, lowest-risk use of the technology in finance, and it is badly underrated.

The numbers come from the data. Always. But "crude is up 3.1%, the dollar is down 0.4%, energy is the only sector green and the majors have not extended" is a sentence that required someone to notice those four facts belong together. Composition, emphasis, and knowing which clause goes first — that is writing, and writing is judgment.

3. Reading unstructured text for meaning, not for keywords

Earnings calls, filings, central bank statements. A keyword scan tells you the word "headwinds" appeared four times. It cannot tell you that management sounded more defensive than last quarter, that a previously firm guidance range quietly became "approximately," or that a risk factor which used to be boilerplate now has three new sentences attached to it.

That is comprehension. It is genuinely hard, genuinely valuable, and there is no regex for it.

4. Classifying things that resist a threshold

Some classification is arithmetic wearing a costume. "Is this stock in an uptrend" is a threshold on a moving average — write the function.

But "is this news item actually about the company, or does it just mention the company" is a judgment. "Is this filing describing a real change in the business or repeating last year's language" is a judgment. The tell is whether two careful analysts could reasonably disagree. If they could, you have found a job for a language model. If they could not, you have found a job for an if statement.

5. Explaining a model's output to a human

Our own research runs on classical statistics — Ward clustering, PCA, ordinary regression, distribution fitting. Not a language model in sight, because those problems have correct answers and well-understood algorithms that produce them.

But the output of a clustering routine is a partition and a silhouette score. Turning that into "the market has spent 37 months in a late-cycle regime characterised by tight policy and calm markets, and the two most likely exits from here are…" is translation work. The algorithm found the structure. The model explains it. Neither could have done the other's job.

6. Adversarial checking of its own kind

One genuinely good use of a language model is auditing text — its own or anyone else's. Does this paragraph claim something the source doesn't support? Does this sentence imply a recommendation we didn't intend? That is reading comprehension applied to prose, which is squarely in scope.

Note the asymmetry: use the model to check claims, use code to check figures.

The pattern that actually works

Everything above collapses into one architecture, and it is the one we run:

Data is computed. Prose is generated. Code guards the boundary.

In our daily brief, every figure — sector moves, yields, crude, gold, crypto, EPS, inflation — comes from verified price and macro data and is quoted as given. The model is explicitly forbidden from originating a number. It decides what matters, what the through-line is, and how to say it. Then deterministic code checks its homework before anything ships: citations are verified against the source pool, named individuals are checked against what the articles actually support, figures are matched against the verified blocks, and a smoke test rejects output that is truncated, refuses, or drifts.

That last layer is the part people skip. A language model in production without deterministic verification around it is not a feature, it is an unbounded liability. The verification is not a sign of distrust — it is what makes the trust earnable. You let the model do the part only judgment can do, and you let code prove the part that can be proven.

The uncomfortable corollary

If you apply this test honestly, a lot of "AI-powered" market tooling turns out to be a language model doing arithmetic badly, wrapped in a subscription.

So it is a fair question to ask of anyone selling you AI, including us: which decisions are the model making, and which numbers is it merely reporting? If the answer is that the model is calculating your indicators, you are paying a premium for unreliability. If the answer is that classical code computes everything checkable and the model is doing the part that genuinely needs judgment — selection, synthesis, explanation — then the technology is being pointed at the thing it is uniquely good at.

We would rather be measured on that question than on how much AI we can claim to use.

Signal over noise.