Service · Backtesting

Backtest review before capital is at risk.

A strong backtest is not enough. StatGazer reviews the data, code, assumptions, timing, universe, and validation design behind a strategy result so your team knows what survives independent scrutiny.

The risk

Backtests fail in boring, expensive ways.

The common failures are rarely dramatic. A feature is joined on the wrong date. A universe excludes dead names. A vendor field is revised after the decision time. Costs are applied after selection. Validation is in-sample, but the chart is good enough to move the conversation forward. Those details can turn a high-Sharpe strategy into an allocation mistake.

Data timing

We inspect as-of dates, report dates, revision histories, feature joins, and any transformation that could put future information into the historical test.

Universe construction

We review membership rules, delistings, survivorship, liquidity filters, corporate actions, and whether the universe could actually have been known at the time.

Validation design

We test whether the result depends on in-sample fit, hand-tuned parameters, favorable regimes, unrealistic costs, or benchmarks that do not match the decision.

What we do

Reproduce, challenge, document.

Backtest review starts with reconstruction. We do not merely read a strategy deck. We try to rebuild the reported result, trace the data transformations, identify where assumptions enter, and write findings that a technical committee can inspect.

Review path

  • Confirm the decision the backtest is supposed to support and the standard of evidence required.
  • Rebuild the result from raw or minimally processed data where access allows.
  • Test for look-ahead, survivorship, data leakage, selection bias, parameter fragility, and cost realism.
  • Run targeted out-of-sample, walk-forward, regime, and sensitivity checks.

Deliverables

  • A findings memo with severity, evidence, and decision impact for each issue.
  • A reproducible notebook or review artifact supporting the main checks.
  • A corrected or reconstructed result where the scope and data make that feasible.
  • A remediation list your team can use before allocation, sizing, or further diligence.

Governance

Useful to PMs, quants, ICs, and risk reviewers.

The output is designed for the person who has to explain the result later. A PM needs to know whether the strategy still deserves attention. A quant lead needs to know what to fix. An IC or risk committee needs a concise record of what was checked and what remains uncertain.

Questions we answer

  • Can the reported performance be regenerated from the data and code?
  • What part of the result disappears under point-in-time reconstruction?
  • Which assumptions matter most to the investment conclusion?
  • What evidence is missing before the strategy can be defended?

Boundary

The governance value is a review trail: what was rebuilt, what changed, what could not be tested, and what remediation would matter before allocation. StatGazer does not provide regulatory assurance, audit opinions, or investment recommendations.

Review inputs

The fastest review starts with a narrow question.

A backtest review can range from a focused leakage check to a full reproduction from raw data. The right scope depends on the decision: whether the team is allocating, deciding whether to continue research, reviewing a manager, or preparing a technical committee memo.

When the review is time-sensitive, we sequence the work around the highest-impact failure modes first. That means timestamp discipline, universe construction, costs, validation design, and whether the headline chart can be regenerated without private manual steps.

If the question is broader than one strategy result, start with model validation; that scope covers methodology, monitoring, assumptions, and model-use governance beyond the backtest itself.

What we ask first

What is the strategy claiming, what data produced the result, what changed in live or paper performance, and what review event is coming up? Those answers determine the minimum evidence needed.

What not to send yet

Do not send source code, credentials, holdings, client names, or raw confidential data through the public form. If there is a fit, we put an NDA and access path in place before sensitive material is reviewed.

How remediation is handled

We do not stop at “the backtest is wrong.” The memo separates errors that invalidate the result from issues that can be repaired, then ranks fixes by impact on the investment conclusion.

Professional boundary. Technical model review, validation, research, and engineering consulting — not financial-statement audit, regulatory assurance, investment, legal, or tax advice.

Related note

Read why reported Sharpe is not enough.

The note explains why a headline performance number needs reconstruction before it becomes decision evidence.

Read note

Delivery model

Backtest review is founder-led.

The same senior reviewer scopes the question, reconstructs the evidence, writes the findings, and handles the technical handoff.

See founder credentials

Next step

Send the high-level backtest question.

Describe the strategy, the decision deadline, and the main concern. Keep confidential data out of the form until an NDA and written scope are in place.

Scope this review