# Prepare your files (/docs/creators/files)

What agents receive is only as good as the files you upload.

## What you can upload [#what-you-can-upload]

| Kind                    | Formats                                           |
| ----------------------- | ------------------------------------------------- |
| Documents               | PDF, Word, OpenDocument text, EPUB                |
| Slides and spreadsheets | PowerPoint, Excel, OpenDocument slides and sheets |
| Text and data           | Markdown, plain text, HTML, CSV, JSON             |
| Email                   | .eml files (Outlook .msg isn't accepted)          |
| Images                  | PNG, JPEG, WEBP, GIF                              |

Each file can be up to 10 MB, a PDF up to 2,000 pages, and a dataset can hold up to 50
files. If a file is too big, split it or save a smaller copy, such as a PDF without
embedded images.

## How quarry reads each kind of file [#how-quarry-reads-each-kind-of-file]

quarry turns every file into text, then cuts that text into passages: short excerpts
that keep their file name, section, and page.

* PDFs with real text are read with their layout, so headings, lists, and tables stay
  in order.
* Scanned PDFs and images go through text recognition, which turns a picture of a page
  into text. An image with no text in it gets a short written description instead.
* Word, PowerPoint, Excel, OpenDocument, EPUB, and email files are read the same way as
  PDFs, slide by slide or sheet by sheet.
* CSV and HTML files become plain text. A CSV's column names appear only at the top,
  not in every passage.

quarry skips sections whose whole heading is "Contents", "Index", "Glossary",
"Preface", "Copyright", or the like, and sections that read like a table of contents:
short lines ending in page numbers, or dot leaders. A section called "Index funds" or
"Copyright law" is kept, and so is any section with real sentences. Reference lists
are kept too.

## What buyers see of your files [#what-buyers-see-of-your-files]

Buyers never get a whole file. For each question they pay for, they get a few
passages, each with its file name, section, and page, so they can cite you. Before
paying, agents can also see a short [free preview](/docs/creators/pricing#free-preview-in-search).

Buyers may use the passages they get and quote them with a citation. The
[terms](/terms) don't let them resell passages, republish them in bulk, or use
repeated questions to rebuild your dataset.

File names are public, so give yours the name you'd want on a citation:
"2025-q3-market-outlook.pdf" rather than "draft-v3-FINAL.pdf".

## Tips for clean passages [#tips-for-clean-passages]

* Upload the original digital file rather than a scan of it. Text recognition works,
  but it makes more mistakes than reading real text.
* Keep clear headings. They become the sections agents cite, and the Test tab uses
  them to suggest questions.
* Put files on the same subject in one dataset, and start a new dataset for a new
  subject. Buyers pay one price to search everything in a dataset.
* Remove drafts, duplicates, and pages you don't want to sell before you upload.

## Rights and privacy [#rights-and-privacy]

Sell only what you own or are allowed to sell, such as your own reports, research,
and data, and no one else's personal data unless you have a legal basis to sell
access to it. You keep ownership of everything you upload, and a dataset that breaks
the [terms](/terms) can be taken down.

Only you can see a dataset before you publish it. To prepare your files, their text
is sent to the reading and search providers listed in the [privacy policy](/privacy),
even before you publish. Your files are never used to train models.

Once your files are ready, [check what agents will receive](/docs/creators/review).
