Transparency
Where the data comes from
A research product that will not say where its inputs come from is asking you to trust a black box twice over. Here is every source, what it is used for, and what we are permitted to redistribute.
The sources
ClinicalTrials.gov (API v2)
public domainTrial registry: phase, design, enrollment, endpoints, completion dates. The backbone of the catalyst universe.
US government work, public domain. Commercial reuse permitted with attribution.
openFDA / Drugs@FDA
public domainApproval history, review designations, application and submission records. Used to resolve outcomes and build sponsor track-record features.
US government work, public domain. Commercial reuse permitted.
Federal Register
public domainAdvisory committee meeting notices — the one forward-looking FDA-side source that does not depend on sponsor disclosure.
US government work, public domain.
SEC EDGAR
public domain8-K disclosures announcing action dates, readout guidance and rejections; XBRL financials for cash and runway; Form 4 insider activity; shelf and ATM filings.
US government work, public domain. Access conditions are technical: a declared User-Agent and a request-rate ceiling.
Equity price history (licensed vendor)
licensedSplit-adjusted daily bars including delisted tickers, used to build the event-reaction labels the model is trained against.
Licensed. Derived metrics only.
Exchange short interest
licensedShort interest as a share of float, published twice monthly with a settlement lag.
Exchange-published; redistribution limited.
The line we hold on licensed data
Everything the FDA, the SEC and the NIH publish is a US government work in the public domain, and we can and do republish it. Price history is different: it is licensed, exchange-derived, and the vendor agreement forbids passing the raw data on.
What that means in practice:
- We show derived metrics computed on top of licensed input, which is what a licence permits — never the raw feed itself.
- We buy end-of-day rather than real-time. Forecasting events days to months out does not need a real-time feed, and designing for end-of-day avoids per-user exchange fees entirely.
- Every stored record carries its source, the time we fetched it, and the terms it was obtained under — so we can demonstrate, field by field, that a number came from a public filing rather than from somewhere it should not have.
What we do not do
We do not scrape a competitor's calendar. A compiled calendar can carry database rights, a site's terms of use are a contract, and bulk extraction creates both a legal problem and a dependency on someone who can cut you off. Our forward calendar is assembled from sponsor disclosure and the Federal Register, which is slower to build and is actually ours.
Stale data is labelled, never carried forward
When a feed fails, affected metrics are shown with their vintage and marked stale rather than silently reused. A forecast built on a trial record that changed last week, presented as current, is worse than no forecast at all.