FilingSift benchmark v0.1.0 methodology and limitations
The benchmark corpus is frozen at 25 issuers and 100 filings across technology, banking, industrial, retail, and distressed-company formats. The parser processed all 100 cases, but zero cases have qualified human annotations. The benchmark is blocked and no accuracy claim is allowed.
Dataset design
The corpus includes easy, difficult, and adversarial filing examples. Ground truth must be created through qualified human review of the cited filing, not copied from machine output. Every discovered failure becomes a regression case.
Launch thresholds
- At least 99.5% fidelity for returned numeric value, sign, unit, currency, and provenance.
- At least 95% event precision, with recall reported rather than hidden.
- Zero critical errors in numeric, sign, unit, currency, or provenance dimensions.
- Correct refusal when evidence is missing, contradictory, or uncertain.
- All 100 cases reviewed and qualified signoff recorded.
Current result
Cases scored: 0. Returned-fact fidelity: unavailable. Event precision and recall: unavailable. Qualified signoff: pending. Machine drafts are preparation material and are explicitly excluded from the golden dataset.
What may be claimed
FilingSift may describe the corpus, deterministic checks, source-citation contract, and current machine-checked records. It may not claim benchmarked accuracy, “verified” output, institutional-grade reliability, or the 99.5% target as an achieved result.
Content snapshot: . Current API data may be newer.