FinRisk

Financial intelligence system
bundled samplev0.3.4research system

Structured financial reasoning

Financial risk intelligence that shows its work.

Turn an annual report into a traceable financial risk assessment. FinRisk computes financial metrics, applies structured checks and shows the evidence behind every conclusion.

INPUTCompany filing
PROCESSMetrics · rules · evidence analysis
OUTPUTAuditable risk assessment
2,000company-disjoint E4 cohort
674deterministically verified outcomes
+0.030paired AUROC improvement
0.0015Holm-adjusted p-value

Why FinRisk?

More control than a report uploaded to a chatbot.

FinRisk separates calculation, interpretation and verification so the reasoning remains inspectable—and can stop when the evidence is not good enough.

01

Traceable

Material conclusions link back through facts, metrics and rules to supporting evidence.

02

Failure-aware

Missing, weak or conflicting evidence can trigger review or abstention instead of a guess.

03

Hybrid by design

Deterministic components own the numbers and core logic; language models are constrained to interpretation.

How the system works

A controlled path from filing to decision.

The Agent orchestrates the work. It does not own the numbers, thresholds or final truth. Each layer has one explicit responsibility and one inspectable output.

01

Ingest

Annual filings, XBRL facts and page-aware documents.

02

Compute

Ratios, trends and traditional models run deterministically.

03

Interpret

Constrained language models extract claims, never final scores.

04

Verify

Every material conclusion must resolve to source evidence.

05

Decide

Failure-aware fusion can flag, review, pass or abstain.

CORE CONTROL

No verified evidence path → REVIEW or ABSTAIN

evidence → fact → metric → rule/model → dimension → decision

External validation · E4

Measured on unseen companies, not polished examples.

The locked v0.3.4 architecture was evaluated out of time on a company-disjoint SEC cohort. The positive result belongs to temporal structured signal—not to an unqualified claim about AI or default prediction.

PRIMARY COMPARISON · VERIFIED SUBSET

Temporal structure added ranking signal

ESTABLISHED_E4
B0 · Ratios only0.678
B6 · Temporal risk0.708
PAIRED ΔAUROC+0.030
95% CI+0.014 — +0.048
ADJUSTED P0.0015

In plain language: adding multi-period financial signals improved how the system ranked future financial-deterioration risk among unseen companies.

Explore a frozen assessment

One assessment. Every layer visible.

Use the tabs to inspect a bundled synthetic result, its evidence chain and decision paths. Nothing here is generated live; the example is frozen and UNCALIBRATED.

STATIC DEMO

Bundled offline sample. The assessment is a real pipeline output for the repository's synthetic fixture - not a real issuer, and not a live run. Reliability is UNCALIBRATED.

Risk index · heuristic, not a probability
95/100

Critical

DecisionABSTAIN

Northstar Components (Synthetic) · FY2025

UNCALIBRATED — not a probability

Severitycritical
Trajectoryinsufficient_history
Evidence coverage0.045
Evidence quality0.900
Model disagreement0.435
Verified paths2/17

The risk index above is the fusion score the decision was derived from. The earlier expert-weighted aggregate for the same evidence was 53; it is retained as a diagnostic only.

Decision basis

Why this decision

degraded run

The system declined to decide because the effective evidence paths did not meet the policy proof floor. Evidence coverage is 0.045.

Reason codes

  • INSUFFICIENT_EVIDENCE

    Evidence coverage fell below the policy floor. The system abstains rather than guessing.

  • CRITICAL_DIMENSION_ESCALATION

    A single dimension reached the critical band and escalated the aggregate as a monotonic floor.

  • UNVALIDATED_RELIABILITY

    Reliability is UNCALIBRATED. The evidence-quality index is not a probability and is not treated as one.

Failure states

  • conflicting_evidence

    Narrative and numeric evidence conflict on at least one verified claim.

Analyst · Critic · Verifier

Verifier verdict REJECTED after checking 16 claim(s), with 15 blocking challenge(s).

What is missing

  • free_cash_flow_growth: N/A — Denominator is zero. Impact: assessment confidence reduced.

Evidence quality composition

  • core data completeness1.00
  • numeric provenance coverage1.00
  • verified claim coverage1.00
  • applicable model coverage0.00
  • temporal depth1.00

These components combine into the evidence-quality index. None of them is a probability that the decision is correct.

System controls

Designed for challenge, not blind trust.

Severity, evidence quality, coverage, disagreement and reliability remain separate quantities throughout the interface.

68

Versioned rules

Thresholds live in inspectable policy, not scattered application code.

04

Traditional models

Altman, Beneish, Piotroski and Ohlson run with applicability checks.

08

Risk dimensions

Coverage and missingness stay visible beside every dimension score.

100%

Replayable

Inputs, versions and material evidence paths are retained for audit.