Documents in. A CSV you can check, line by line.

Send us your PDFs, scans, or images and the list of fields you need back. You get a spreadsheet where each value either carries the file it came from and the page number it was read on, or is marked as one we could not tie to a page — so you can verify our work instead of trusting it.

$245, fixed

Up to 250 documents. Up to 20 fields. Five business days from the day we agree the field list.

Invoiced by CyberNative AI LLC. No account to create, no card on file, no subscription.

What arrives

  • data.csv and data.xlsx — one row per document, one column per field you asked for.
  • A source column for every field — the file name and the page number each value was read from.
  • An evidence column for every field, with one of three plain values: cited (we found a value and can point at the page it is printed on), uncited (we found a value but cannot point at a page), or not_found (we found nothing and left the cell blank). It is a fact about the value, not a score we invented.
  • flagged.csv — every row that needs a human look, with the specific reason: a field we could not find, or a value we could not tie to a page. A value we could not find stays blank. Nothing is guessed to make the file look full.
  • reconciliation.md — documents in against rows out, cells filled against cells blank, and the fill rate for each field. Any document that produced nothing is named.
  • data-dictionary.md — how each field was normalised: date format, whitespace, casing.

How to check our work in ten minutes

Open the CSV. Pick five rows at random. Open the file each one names, at the page it names, and read the value off the document yourself.

That is the whole acceptance test. If a value marked cited is not the correct value for that field on the page it cites, we fix it free or refund you in full.

Note what that promise is measured against: the correct value for the field, not merely a string that appears somewhere on the page. Those are not the same test, and the difference is the subject of the accuracy paragraph below.

Two things it deliberately does not cover, said here rather than in small print:

  • A blank cell cites no page, so it cannot fail that test. Blanks get their own promise instead — the coverage one below.
  • A value marked uncited is one we could not tie to a specific page, so you cannot run the ten-minute check on it. It is marked precisely so you know which values those are.

The coverage promise. Your quote names the fill rate we expect for each field, measured on documents you send us before you pay anything. If a field arrives filled at less than half its quoted rate, that is the same deal — we fix it free or refund you in full. We would rather quote you a low number, or turn the job down, than hand you a spreadsheet of empty cells.

What we have actually measured, including the parts that look bad

We are a new seller with no customers yet. Here is every number we have, from our own runs on public filings from the US SEC. There are three, and the newest one is the worst. Read all three before you decide anything.

Run one — 40 filings, six fields, the cover page of each. We took 40 periodic reports from 40 different companies and asked for six fields that live on a filing's cover page: issuer name, form type, period end, SEC file number, state of incorporation, and address of principal executive offices. We pointed the extraction at the cover page of each document. All six fields came back filled on all 40 documents — 40 out of 40, six times over. 239 of those 240 values carry the page they were read from. One does not, and it is marked uncited rather than quietly dropped.

Run two — 100 documents, eight fields, an earlier and simpler setup, and 73 of them came back empty. Before run one we did a wider pass: eight fields across 100 documents, reading the first three pages of each document as plain text rather than a page picked out for cover markers, using a plain rules-based extractor rather than the setup a job would use today. The eight were the six above plus filing date and CIK number. 73 of the 100 rows came back completely blank. Four rows came back with all six of the fields above. 96 of the 100 rows were flagged for a human to look at, each with a reason. The two extra fields, filing date and CIK, came back empty on all 100 documents, and we dropped them from the field list before run one. So the six fields in run one are a list we narrowed after two of eight failed outright — read the 40 out of 40 above with that in mind.

We are showing you a result from an older setup because it is the widest number we have, and because of what it tells you that run one cannot: when a field is not on the pages we read, you get an empty cell and a flagged row. Not a plausible-looking guess.

Be clear about where that failure sits. 98 of those 100 filings run longer than three pages, so a field printed past page three was never looked at. Some part of those 73 blanks is a limit of how we ran it, not proof the field was missing from the document. We have not measured how large that part is, and we are not going to guess at it here.

Run three — 25 filings, the same six fields, the same reader, and not one clean row. This is the newest run we have and it is the worst one. We took 25 SEC filings and asked for the same six cover-page fields as run one, using the same vision-based reader run one used. 59 of the 150 cells came back filled, and not one of the 25 rows came through clean — all 25 were flagged for a human. Issuer name came back on 22 of the 25 documents; SEC file number on 3. Three rows did carry all six fields, and even those were flagged. The difference between this and run one that we can actually point at in our own records is what each run was handed: run one got exactly one page per document, picked out by us because it carried SEC cover-page markers, and run three's record shows no such page being picked. Read the 40 out of 40 above as a number that depends on choosing the page first.

And the number we are not going to give you: accuracy. We hand-checked 20 values per field from run one against an outside record, and afterwards found the checker itself was unreliable — it had been written while its own answers were visible, so it flagged some correct values as wrong and would have waved some wrong values through. We threw the result out rather than quote the flattering half of it. So we have no accuracy percentage we would stand behind, and we are not going to publish one until we have a check that was frozen before it saw a single answer. If you see an extraction vendor quoting you a headline accuracy number, ask them what it was measured against and who wrote the comparison.

Those four paragraphs are the whole evidence base. What they are good for: run one says that on 40 documents with one cover page picked out for each, all 240 cells came back filled and 239 of those values carried a page citation — it does not say those values were correct, because the only check we ran on that is the one we threw out. Run two says that when the field is not on the pages we read, you get a blank and a flag instead of a guess. Run three says the same reader, on the same six fields and without a cover page picked out for it, returned 59 of 150 cells and no clean row at all. None of the three tells you what will happen to your documents — which is exactly why the next section exists and why it is free.

Not sure your fields are even there? Find out before you pay anyone.

Send two or three of your documents and the list of fields you want. We run them and send you back — for free — the fill rate we got for each field.

It is the same number, measured the same way, that arrives with the finished job, and it is the number your quote is held to. The answer is sometimes "field X is not on these pages", and that answer is worth having before you pay anyone, including us.

Where your files go

Onto hardware we own, into a container with outbound networking switched off. No document, page image, or extracted value is sent to a third-party model API.

What we do not do

We never touch your systems, log into anything of yours, or hold a credential of yours. You send files; we send a spreadsheet back. We do not scrape websites.

Order it

Email hello@cybernative.ai with:

  1. two or three example documents, and
  2. the list of fields you want.

You get back a fixed quote, and an invoice from CyberNative AI LLC if the job fits. If it does not fit the scope above, we will say so and tell you what it would take instead.