Service · Independent review

Independent model validation.

StatGazer is an independent model validation consultancy for investment teams that need a second technical read on forecasting, backtesting, risk, and machine-learning models before capital, credibility, or governance depends on the output.

When it matters

Validation creates a taxonomy of findings.

A model can look strong in research and still hide different classes of risk: unavailable features, weak conceptual assumptions, brittle calibration, undocumented monitoring, or evidence that cannot be reproduced. A useful validation separates those issues instead of collapsing them into a generic pass/fail view.

Methodology review

We read the model as a technical system: assumptions, target definition, feature construction, validation design, calibration, and the link between statistical evidence and the business decision.

Leakage and timing

We test whether inputs were available at the claimed decision time, whether joins are point-in-time, and whether universe construction, survivorship, or revision history contaminates the result.

Reproducibility

We trace the path from source data to output. The goal is not a polished slide; it is evidence your team can regenerate, inspect, and defend under technical review.

How we work

Independent challenge, written evidence.

The review follows a disciplined model-risk pattern: conceptual soundness, outcomes evidence, monitoring, limitations, and documentation. SR 11-7 is useful vocabulary for that discipline, used here as a practical parallel rather than a claim of regulatory assurance.

What we examine

  • Data lineage, timestamp discipline, joins, missingness, and revision history.
  • Feature construction, target definition, model choice, assumptions, and parameter stability.
  • Validation design: train/test split, walk-forward testing, regime sensitivity, calibration, and benchmark comparisons.
  • Operational controls: versioning, monitoring, documentation, and handoff readiness.

Finding taxonomy

  • Confirmed defects with evidence and decision impact.
  • Material assumptions that are defensible only under stated conditions.
  • Unresolved questions caused by missing evidence, access, or documentation.
  • Remediation items ranked by model-use risk, not cosmetic neatness.

Commercial fit

Built for the buyer who has to defend the answer.

The strongest use case is not “we need a consultant.” It is “someone will ask whether this model is real, and our internal team needs a clean technical record before that conversation.”

Good fit

  • A strategy backtest is moving toward allocation or increased sizing.
  • A risk or forecasting model is being used in a recurring decision process.
  • A PM, CIO, risk committee, IC, or vendor-review process needs defensible evidence.
  • The internal team wants an outside reviewer without handing away model ownership.

Not a fit

  • You need a financial-statement audit opinion or regulatory attestation report.
  • You want someone to bless a model without access to evidence.
  • You need investment advice, legal advice, tax advice, or a performance guarantee.
  • You want confidential data reviewed before an NDA and scope are in place.

Working inputs

Validation starts with context, not a file dump.

Before we request code or data, we define the model use, the decision it supports, the review audience, and the level of evidence that would change the decision. That keeps the engagement focused and prevents a broad technical fishing expedition.

If the model is still early, the right engagement may be a narrower methodology review. If the model is already in production or near an allocation decision, the scope usually needs stronger evidence: reproducibility, out-of-sample behavior, failure modes, and a remediation path your team can execute. The point is to make the review decision-useful, not merely comprehensive.

What helps upfront

A short model overview, the business decision, current validation evidence, known concerns, and the deadline or review event the work has to support. Confidential artifacts can wait until NDA and scope are agreed.

Evidence package

The review package ties each finding to supporting artifacts: code paths, data checks, validation outputs, screenshots, or documentation gaps. That keeps the memo inspectable after the readout.

Handoff standard

The final readout is written for both technical owners and decision makers. Your team should leave knowing what was checked, what changed the conclusion, what to fix first, and what should be monitored after deployment.

Narrower question

If the issue is one strategy backtest, start narrower.

Backtest review focuses on reconstruction, leakage, survivorship, costs, and validation design around a specific strategy result.

See backtest review

Reference asset

Use the public review standard as the checklist.

The review standard explains the evidence structure behind model validation: decision context, data timing, reproduction, validation design, finding taxonomy, and handoff boundary.

Open review standard

Delivery model

Model validation is founder-led.

Scoping, technical review, findings, and handoff are handled directly by Evgenii Azarov, PhD, rather than passed through a junior delivery bench.

See founder credentials

Related note

Classify findings before debating severity.

The finding taxonomy note explains how to separate confirmed defects, material assumptions, unresolved questions, limitations, and remediation items.

Read finding taxonomy

Professional boundary. Technical model review, validation, research, and engineering consulting — not financial-statement audit, regulatory assurance, investment, legal, or tax advice.

Next step

Start with the model and the decision it supports.

Send a high-level description. Do not include confidential data, code, credentials, portfolio holdings, or client names until an NDA and written scope are in place.

Request a scoping call