I maintain a multi-agent stock research pipeline: four analyst agents run concurrently, a bull/bear debate engine argues over their reports, a research manager adjudicates, and a synthesizer merges everything into a briefing with a signal and a conviction score.
For a long time it had one embarrassing property. Every run was a cold start.
The pipeline already produced the raw material for a track record. Backtest trials carried a realized_return. The scorer computed hit rates. But nothing fed any of it back. The same wrong call on the same ticker could be made indefinitely and the system would never notice.
Closing that loop is easy to describe and easy to get subtly, silently wrong. This post is about the two decisions that made it safe.
The naive version
The obvious design: append every resolved call to a per-ticker log, and inject the recent history into the synthesis prompt.
1class OutcomeRecord(BaseModel):
2 ticker: str
3 as_of_date: date # when the call was made
4 horizon_days: int
5 signal: Signal
6 conviction_score: float
7 signal_convergence: float
8 entry_price: float | None = None
9 exit_date: date | None = None
10 exit_price: float | None = None
11 realized_return: float | None = None
12 source: str = "backtest"
13Storage is append-only JSONL under data/<TICKER>/outcomes.jsonl, deduplicated by (as_of_date, horizon_days, source). That part is boring and correct.
The interesting part is the read. In live use you want everything. In a backtest you must only see what was knowable at the time — and this is where the naive filter shows up:
1# WRONG
2def visible_on(self, as_of: date | None) -> bool:
3 if as_of is None:
4 return True
5 return self.as_of_date < as_of # "the call was made in the past, so it's fine"
6It reads like the right thing. The record was created before the analysis date, so surely it's historical.
It isn't.
Why gating on entry date destroys the backtest
Consider a trial dated 2024-01-01. A prior record has:
as_of_date = 2023-12-01— the call was made a month beforeexit_date = 2024-05-01— the 150-day horizon resolves five months after
The entry-date filter admits it. The record then hands the synthesizer a realized_return measured on 2024-05-01, while the model is pretending it's January. The prompt is now literally telling the model what the stock did over the next four months.
The position was opened in the past. Its outcome had not happened yet. Those are different facts, and only the second one matters.
What makes this bug nasty is that it doesn't crash, doesn't warn, and doesn't look wrong in the output. The backtest still runs. The report still prints a hit rate. That hit rate is now partly a measurement of how well the pipeline can read a number it was handed. Every downstream metric — directional Sharpe, information coefficient, the conviction-vs-return correlation — inherits the contamination. You would ship it, and it would look better than before.
The correct guard is one line different and gates on the exit:
1def visible_on(self, as_of: date | None) -> bool:
2 """Whether this record is knowable when analyzing `as_of`.
3
4 Gated on exit date: a trade entered before `as_of` but closed after it
5 has an outcome that had not happened yet.
6 """
7 if as_of is None:
8 return True
9 if self.exit_date is None:
10 return False
11 return self.exit_date < as_of
12Three details in there are deliberate:
exit_date is None→ hidden. An unresolved record has no outcome to learn from. It contributes nothing to a dated run and is excluded rather than treated as a pending zero.- Strictly less than. A record exiting on the analysis date is hidden. Same-day resolution is a coin flip on intraday timing, and the cost of being wrong here is silent contamination, so it loses.
as_of is Nonemeans live. A real-time run has no future to leak from and sees the full history.
The store applies the filter at load time, so no caller can forget it:
def load(self, ticker: str, before: date | None = None) -> list[OutcomeRecord]:
...
records = [r for r in records if r.visible_on(before)]
return sorted(records, key=lambda r: (r.as_of_date, r.horizon_days))
And the orchestrator passes the trial's date straight through, so a backtested run and a live run take the same code path with a different argument:
1if self.settings.enable_outcome_memory:
2 outcome_store = OutcomeStore(self.settings.data_dir)
3 prior = outcome_store.load(ticker, before=self.as_of_date)
4 if prior:
5 memory_context = build_memory_context(
6 prior, outcome_store.calibration(ticker, before=self.as_of_date)
7 )
8There is exactly one place that reads history without the filter — save_calibration(), which dumps a full-history summary to calibration.json for dashboards. Its docstring says why in as many words: this file is for humans looking back, never an input to a dated analysis. An unfiltered read that is easy to mistake for a filtered one deserves a comment explaining that it is intentional.
The leakage tests are the ones I'd point a reviewer at first, because they encode the trap rather than the happy path:
1def test_entry_in_the_past_does_not_make_a_future_exit_visible(self):
2 # The trap this guard exists for: entry is old, but the outcome is not
3 # yet known on the analysis date.
4 record = _record("2023-12-01", "2024-05-01")
5 self.assertFalse(record.visible_on(date(2024, 1, 1)))
6
7def test_exit_on_the_analysis_date_itself_is_hidden(self):
8 record = _record("2024-01-01", "2024-03-01")
9 self.assertFalse(record.visible_on(date(2024, 3, 1)))
10The second decision: calibration is reported, never applied
Once you have a track record, the tempting next step writes itself. You know the hit rate. You know whether high-conviction calls beat low-conviction ones. So scale the conviction score by it — a ticker where the system has been 40% right should produce weaker signals.
Nothing in the module does this. compute_calibration() returns a CalibrationSummary and that summary is rendered into prose. It never touches a number in the briefing.
1class CalibrationSummary(BaseModel):
2 """Deterministic scoring of a ticker's resolved history."""
3
4 ticker: str
5 trials: int = 0
6 directional_trials: int = 0
7 neutral_trials: int = 0
8 hit_rate: float | None = None
9 mean_return: float | None = None
10 high_conviction_hit_rate: float | None = None
11 low_conviction_hit_rate: float | None = None
12 conviction_separates: bool | None = None
13 ...
14
15 @property
16 def sufficient_sample(self) -> bool:
17 return self.directional_trials >= MIN_TRIALS_FOR_SIGNAL
18The argument for restraint is a sample-size argument. A ticker with fifteen resolved directional calls, spread over a couple of years and a few different market regimes, has a hit rate with an enormous confidence interval. Nine hits out of fifteen is 60%; ten is 67%. One trade moves the "coefficient" by seven points. Multiply live signals by that and you have built a system that overfits to a handful of historical accidents — and it will do it invisibly, because a rescaled conviction score looks exactly like an honest one.
That would be a strictly worse failure than the cold start it replaced. A cold start is obviously ignorant. A confidently miscalibrated multiplier is ignorance with a decimal point on it.
So the threshold exists only to change the wording, never the math:
# Below this, a hit rate is noise. Reported with the count attached rather than
# suppressed, but never described as a track record.
MIN_TRIALS_FOR_SIGNAL = 8
Under eight directional trials, the prompt fragment appends a caveat instead of hiding the number:
1if calibration.hit_rate is not None:
2 confidence_note = (
3 "" if calibration.sufficient_sample else " — too few trials to be a track record"
4 )
5 lines.append(f"Directional hit rate: {calibration.hit_rate:.0%}{confidence_note}.")
6The same restraint runs through the rest of the arithmetic:
- Neutral calls are excluded from the hit rate.
correctreturnsNonefor them. A neutral call has no direction to be right about, so scoring it either way would be inventing an opinion the pipeline explicitly declined to have. Neutral is a valid output in this system, and it's counted separately. - Unresolved records are ignored, not counted as misses. Missing data is missing, not evidence of failure.
conviction_separatesstaysNoneon a tiny sample. It's only computed when both the high- and low-conviction groups have at least three trials — and even then it's a boolean shown to the reader, not a weight.
The injected prompt fragment ends by telling the model, in plain terms, what the data is and isn't:
How to use this: it is evidence about this system's past accuracy on this ticker, not evidence about the stock's future. Do not flip a view the current data supports merely because past calls missed, and do not extrapolate a small sample. If prior calls failed for a reason that is still present in today's data, say so explicitly in your uncertainties.
One more small thing I like: with no history at all, build_memory_context() returns None rather than a "no track record available" paragraph. Telling a model it has no track record invites it to write a sentence about the absence of a track record. Silence is cheaper.
Keeping the backtest reproducible
Recording outcomes is opt-in — stock-analysis-backtest --record-outcomes, off by default. Writing outcomes on every backtest would mean a second run over the same date range reads the first run's results, and two runs of the same experiment would stop being comparable. The feedback loop has to be something you switch on deliberately, not a side effect of measuring.
This sits alongside a related guard in the scorer, which partitions trials against the models' approximate training cutoffs and reports the post-cutoff slice separately as the clean out-of-sample test. Different leak, same discipline: if the model could have known the answer, say so out loud rather than averaging it into the headline number.
What generalizes
Two things I'd carry to any system that learns from its own history.
Ask what the timestamp actually means. A record has several dates on it and only one of them answers "when did this become knowable?" Picking the wrong one produces a system that passes its tests, reports good numbers, and is measuring nothing.
A statistic you can compute is not a coefficient you should apply. The gap between "here is the hit rate, with the trial count attached" and "conviction × hit rate" is the entire difference between reporting and overfitting. Surfacing the number to a human reader costs nothing if it's wrong. Wiring it into the output costs you the ability to tell that it's wrong.