Structured financial reasoning
Financial risk intelligence that shows its work.
Turn an annual report into a traceable financial risk assessment. FinRisk computes financial metrics, applies structured checks and shows the evidence behind every conclusion.
Why FinRisk?
More control than a report uploaded to a chatbot.
FinRisk separates calculation, interpretation and verification so the reasoning remains inspectable—and can stop when the evidence is not good enough.
Traceable
Material conclusions link back through facts, metrics and rules to supporting evidence.
Failure-aware
Missing, weak or conflicting evidence can trigger review or abstention instead of a guess.
Hybrid by design
Deterministic components own the numbers and core logic; language models are constrained to interpretation.
How the system works
A controlled path from filing to decision.
The Agent orchestrates the work. It does not own the numbers, thresholds or final truth. Each layer has one explicit responsibility and one inspectable output.
Ingest
Annual filings, XBRL facts and page-aware documents.
Compute
Ratios, trends and traditional models run deterministically.
Interpret
Constrained language models extract claims, never final scores.
Verify
Every material conclusion must resolve to source evidence.
Decide
Failure-aware fusion can flag, review, pass or abstain.
No verified evidence path → REVIEW or ABSTAIN
evidence → fact → metric → rule/model → dimension → decisionExternal validation · E4
Measured on unseen companies, not polished examples.
The locked v0.3.4 architecture was evaluated out of time on a company-disjoint SEC cohort. The positive result belongs to temporal structured signal—not to an unqualified claim about AI or default prediction.
Temporal structure added ranking signal
In plain language: adding multi-period financial signals improved how the system ranked future financial-deterioration risk among unseen companies.
Explore a frozen assessment
One assessment. Every layer visible.
Use the tabs to inspect a bundled synthetic result, its evidence chain and decision paths. Nothing here is generated live; the example is frozen and UNCALIBRATED.
Bundled offline sample. The assessment is a real pipeline output for the repository's synthetic fixture - not a real issuer, and not a live run. Reliability is UNCALIBRATED.
Critical
Northstar Components (Synthetic) · FY2025
UNCALIBRATED — not a probability
The risk index above is the fusion score the decision was derived from. The earlier expert-weighted aggregate for the same evidence was 53; it is retained as a diagnostic only.
Decision basis
Why this decision
degraded run
The system declined to decide because the effective evidence paths did not meet the policy proof floor. Evidence coverage is 0.045.
Reason codes
INSUFFICIENT_EVIDENCEEvidence coverage fell below the policy floor. The system abstains rather than guessing.
CRITICAL_DIMENSION_ESCALATIONA single dimension reached the critical band and escalated the aggregate as a monotonic floor.
UNVALIDATED_RELIABILITYReliability is UNCALIBRATED. The evidence-quality index is not a probability and is not treated as one.
Failure states
conflicting_evidenceNarrative and numeric evidence conflict on at least one verified claim.
Analyst · Critic · Verifier
Verifier verdict REJECTED after checking 16 claim(s), with 15 blocking challenge(s).
What is missing
- free_cash_flow_growth: N/A — Denominator is zero. Impact: assessment confidence reduced.
Evidence quality composition
- core data completeness1.00
- numeric provenance coverage1.00
- verified claim coverage1.00
- applicable model coverage0.00
- temporal depth1.00
These components combine into the evidence-quality index. None of them is a probability that the decision is correct.
System controls
Designed for challenge, not blind trust.
Severity, evidence quality, coverage, disagreement and reliability remain separate quantities throughout the interface.
Versioned rules
Thresholds live in inspectable policy, not scattered application code.
Traditional models
Altman, Beneish, Piotroski and Ohlson run with applicability checks.
Risk dimensions
Coverage and missingness stay visible beside every dimension score.
Replayable
Inputs, versions and material evidence paths are retained for audit.