quarryDocs

Prepare your files

Check which files you can sell, how quarry reads each kind, and how to get clean passages from them.

What agents receive is only as good as the files you upload.

What you can upload

KindFormats
DocumentsPDF, Word, OpenDocument text, EPUB
Slides and spreadsheetsPowerPoint, Excel, OpenDocument slides and sheets
Text and dataMarkdown, plain text, HTML, CSV, JSON
Email.eml files (Outlook .msg isn't accepted)
ImagesPNG, JPEG, WEBP, GIF

Each file can be up to 10 MB, a PDF up to 2,000 pages, and a dataset can hold up to 50 files. If a file is too big, split it or save a smaller copy, such as a PDF without embedded images.

How quarry reads each kind of file

quarry turns every file into text, then cuts that text into passages: short excerpts that keep their file name, section, and page.

  • PDFs with real text are read with their layout, so headings, lists, and tables stay in order.
  • Scanned PDFs and images go through text recognition, which turns a picture of a page into text. An image with no text in it gets a short written description instead.
  • Word, PowerPoint, Excel, OpenDocument, EPUB, and email files are read the same way as PDFs, slide by slide or sheet by sheet.
  • CSV and HTML files become plain text. A CSV's column names appear only at the top, not in every passage.

quarry skips sections whose whole heading is "Contents", "Index", "Glossary", "Preface", "Copyright", or the like, and sections that read like a table of contents: short lines ending in page numbers, or dot leaders. A section called "Index funds" or "Copyright law" is kept, and so is any section with real sentences. Reference lists are kept too.

What buyers see of your files

Buyers never get a whole file. For each question they pay for, they get a few passages, each with its file name, section, and page, so they can cite you. Before paying, agents can also see a short free preview.

Buyers may use the passages they get and quote them with a citation. The terms don't let them resell passages, republish them in bulk, or use repeated questions to rebuild your dataset.

File names are public, so give yours the name you'd want on a citation: "2025-q3-market-outlook.pdf" rather than "draft-v3-FINAL.pdf".

Tips for clean passages

  • Upload the original digital file rather than a scan of it. Text recognition works, but it makes more mistakes than reading real text.
  • Keep clear headings. They become the sections agents cite, and the Test tab uses them to suggest questions.
  • Put files on the same subject in one dataset, and start a new dataset for a new subject. Buyers pay one price to search everything in a dataset.
  • Remove drafts, duplicates, and pages you don't want to sell before you upload.

Rights and privacy

Sell only what you own or are allowed to sell, such as your own reports, research, and data, and no one else's personal data unless you have a legal basis to sell access to it. You keep ownership of everything you upload, and a dataset that breaks the terms can be taken down.

Only you can see a dataset before you publish it. To prepare your files, their text is sent to the reading and search providers listed in the privacy policy, even before you publish. Your files are never used to train models.

Once your files are ready, check what agents will receive.

On this page

View as Markdown