The Number That Looked Like a Discovery — Field Report Nº 003
We tested a paid trading indicator stack against 16 years of audited institutional data. One number came back at 6.60 — a result that would normally be called a discovery. Here's why it was actually worth 2.68, and how we knew before we got excited.
2026-08-11

The Number That Looked Like a Discovery
First, a correction to a promise.
At the end of Field Report Nº 002 we said Nº 003 would cover what happened to our two surviving strategies once they went live. It isn't ready. Both are still running, and the honest position is that a few weeks of live data tells you almost as little as the one day we had when we published. Reporting on it now would be exactly the kind of thing this series exists to avoid. It's coming, as Nº 004, when the tracking period is long enough to mean something.
In the meantime, we tested something else. And it produced the single most instructive number we've generated all year — instructive precisely because it was wrong.
What we tested
Not a robot this time. A bias tool: a paid indicator package built on Commitment of Traders data.
Here's the background, because the data itself is genuinely interesting. Every week, the US futures regulator publishes exactly how the big players are positioned in 20-odd major markets — gold, silver, oil, corn, the major currencies. It splits them into commercial hedgers (the mining companies, the farmers, the banks — the people with an actual business in the underlying), large speculators, and small retail traders. It's been published since 1986. It's audited. It's free.
The theory built on it is decades old and genuinely appealing: the commercial hedgers are the smart money. When they're positioned at an extreme, the market is about to turn. Tools in this category scale that positioning to a 0–100 score and tell you when to be bullish or bearish — bias only, never entries; you're expected to time entries with something else.
That's a testable claim. So we tested it.
Why we wrote the test down before running it
This is the part that matters more than the result.
Earlier this year we burned roughly 40 different configurations on a single strategy idea, found one that worked beautifully, and eventually established that it was noise — we'd effectively kept rolling the dice until we got a good number, then called the good number a discovery. That's not fraud, it's the single most common way honest people fool themselves with data.
So now, before touching the data, we write the whole test down: what counts as success, at what threshold, measured how. Then we run it once and we live with the answer.
For this one, fixed in advance:
- Bucket the positioning score into five bands, 0–20 through 80–100
- Measure the average return over the following 1, 4, 8, 13 and 26 weeks
- The bands must line up in order — more bullish positioning, better returns. Not just the extremes working
- The statistical strength must clear 3 (we set it above the usual bar of 2 deliberately, because we were testing several horizons at once)
- The effect must survive the publication delay — the data describes Tuesday's positions but isn't published until Friday, and any test that trades on Tuesday's data is quietly cheating
Nine markets, 16 years, 7,686 weekly observations.
The result
The direction was right. Low commercial positioning did precede weaker returns; high positioning preceded stronger ones, consistently, across most markets and most horizons.
And it failed anyway, on two of the four criteria.
The bands didn't line up in order. At every single horizon we measured, the 60–80 band beat the 80–100 band. Read that again in the context of a tool whose entire premise is extremes: the most extreme institutional positioning performed worse than merely elevated positioning. Five out of five horizons. That's not a weak result, it's the wrong shape — the opposite of what the theory requires.
And then the number. The raw statistical strength on the best horizon came back at 6.60. For context, 2 is the conventional threshold for "probably real" and 3 is a strict bar. 6.60 is the kind of number that makes you start writing the strategy.
It was worth 2.68.
Why 6.60 was really 2.68
Two flaws, both invisible if you don't go looking, both of which inflate results in the flattering direction.
Overlapping windows. We take a reading every week and measure the return over the next 13 weeks. So this week's reading and next week's reading describe almost the same 13 weeks — they overlap by 12. We had 7,686 observations, but nothing like 7,686 independent pieces of evidence. The statistics assumed independence and rewarded us for it.
Correlated markets. We pooled nine markets to get a bigger sample. But when the dollar moves, the euro, the pound, the Aussie and the yen all move together. Nine correlated markets are not nine independent tests — they're closer to two or three. Pooling made the sample look large while the actual information barely grew.
Correct for both — collapse the markets into one combined weekly reading, then apply a standard adjustment for the overlap — and 6.60 becomes 2.68. Below our bar of 3. The best result across every horizon we measured only reached 2.31.
Nothing dishonest happened to produce the 6.60. It's what the standard formula returns on that data. It is simply a number that means something different from what it appears to mean, and there is an entire industry that never applies the correction.
The one thing that passed cleanly
The publication-delay check came back essentially neutral — trading the data on Friday when it's actually published, versus cheating and trading it on Tuesday, differed by hundredths of a percent.
That's worth knowing for its own sake: over weekly-to-quarterly horizons, the three-day delay costs nothing, because institutional positioning shifts far more slowly than three days. It also told us something more useful — the weak result wasn't an artefact of our own caution. The pipeline was clean. The effect really is that small.
What we concluded
We stopped. The tool has three further layers we could have tested — open interest, seasonality, a valuation model — but we'd committed in advance to a rule: if the foundation isn't there, no execution overlay on top of it can rescue it. The foundation isn't there. Adding layers at that point isn't research, it's shopping for a result.
To be fair to the theory: our nine markets were mostly currencies, where "commercial hedgers" are bank trading desks rather than businesses hedging real physical goods. The idea deserves a proper test on corn, wheat, coffee and sugar, where the hedgers are actual farmers and processors. That's a genuinely different test and we'll write it down before we run it, like this one.
What we won't do is present that as a second attempt at making this result work.
Why any of this belongs on a business blog
Because it's the same discipline, and most of our clients are buying the discipline rather than the trading.
Somebody is going to sell you a number this year. Conversion rates, engagement figures, ad performance, a dashboard where everything trends up and to the right. The questions that actually protect you are the same three we used here: What would have counted as failure, and was that decided before or after seeing the data? How many things were tried before this one was shown to me? And is this number measuring what it appears to measure?
We build client systems the same way we test strategies — which is why our reporting tends to show the flat months rather than smoothing them out. A number you can trust when it's bad is the only kind worth having when it's good.
Field Report Nº 004: the live results on the two survivors from Nº 002 — whatever they turn out to be.
Method note: CFTC Commitment of Traders legacy reports 2010–2026, mapped by contract code; daily prices via the MetaTrader 5 API. Nine markets (seven FX pairs, gold, silver), 7,686 market-weeks. Signals joined to price at the Friday publication timestamp. Headline significance computed on an equal-weighted weekly cross-market series with Newey-West HAC standard errors at lag = horizon; the uncorrected pooled figure is reported alongside it for comparison. Test design and success criteria were fixed in writing before the analysis was run. Historical simulation, not investment advice.
Don't just read about it
Need a professional website? We build it.
Mobile-friendly, Google-ready and live in 7 days — built in PMB for SA businesses.
from R2,500 · once-off · live in 7 days
Prefer email? nkanyiso@inkatech.co.za · patricia@inkatech.co.za
Keep reading
More guides for your business

How Much Does a Website Cost in South Africa in 2026?
Explore website costs in South Africa for 2026. Discover pricing, factors affecting cost, and how to get a professional website.
Read article →
How Much Does Social Media Video Cost in South Africa? (2026 Guide)
What short-form video content actually costs in South Africa in 2026 — agency rates, freelancer rates, and what you should expect to pay per clip.
Read article →