A lot of useful data on the web does not live on normal webpages. For example, annual reports, financial statements, research papers, government publications, investor presentations, price lists, and product specifications are often published as PDFs.
The information is public, but using it in an application is not always straightforward because most scraping tools are built around HTML.
A PDF can look perfectly readable in your browser while giving a traditional scraper almost nothing useful. Some documents contain real text, others contain scanned pages, and some websites force the file to download instead of rendering it normally.
Spidra now handles document URLs directly. You can scrape a PDF into clean Markdown, extract specific fields into structured JSON, or start with Search when you do not know the document URL yet. The same document support also covers DOC/DOCX, XLSX, PPT/PPTX, ODT, RTF, EPUB, and CSV files.
That means a document workflow can be as simple as:
PDF URL → Markdownor continue further:
Search → Find PDF → Scrape → Structured JSONScrape a PDF into Markdown
If you already have the PDF URL, you can send it directly to the Scrape API.
curl -X POST https://api.spidra.io/api/scrape \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"urls": [
{
"url": "https://example.com/annual-report.pdf"
}
],
"output": "markdown"
}'Spidra detects that the URL points to a document and converts its contents to Markdown instead of trying to parse it like HTML. Scanned or image-only PDF pages are also transcribed, so they don't come back empty.
This is useful when you want the document itself for something like RAG, summarization, search, archiving, or further processing.
The same request shape works for other supported document formats too, including Word documents, spreadsheets, presentations, EPUB files, and CSVs.
Extract structured JSON from a PDF
Sometimes the document isn't what you really want. The data inside it is.
Suppose you are processing a financial report and only need the company name, reporting period, revenue, net income, and earnings per share. Returning the entire report as Markdown means your application still has to work out where those values are and how to structure them.
With Spidra, you can add a standard JSON Schema to the scrape request and define the shape you want back.
For example:
curl -X POST https://api.spidra.io/api/scrape \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"urls": [
{
"url": "https://example.com/apple-financial-report.pdf"
}
],
"prompt": "Extract the financial data from this report. Keep monetary values as numbers in millions.",
"schema": {
"type": "object",
"required": [
"company",
"reporting_period",
"period_end_date",
"income_statement_three_months_ended"
],
"properties": {
"company": {
"type": "string"
},
"reporting_period": {
"type": "string"
},
"period_end_date": {
"type": "string"
},
"income_statement_three_months_ended": {
"type": "object",
"required": [
"total_net_sales",
"net_income",
"earnings_per_share_diluted"
],
"properties": {
"total_net_sales": {
"type": "object",
"required": [
"current_period",
"prior_year_period"
],
"properties": {
"current_period": {
"type": ["number", "null"]
},
"prior_year_period": {
"type": ["number", "null"]
}
}
},
"net_income": {
"type": "object",
"required": [
"current_period",
"prior_year_period"
],
"properties": {
"current_period": {
"type": ["number", "null"]
},
"prior_year_period": {
"type": ["number", "null"]
}
}
},
"earnings_per_share_diluted": {
"type": "object",
"required": [
"current_period",
"prior_year_period"
],
"properties": {
"current_period": {
"type": ["number", "null"]
},
"prior_year_period": {
"type": ["number", "null"]
}
}
}
}
}
}
}
}'The schema sits at the top level of the scrape request alongside prompt, and adding a schema automatically makes the output JSON. Spidra accepts standard JSON Schema with nested objects, arrays, enums, nullable fields, and other supported schema properties.
The prompt and schema do different jobs. The prompt tells Spidra how to interpret the document, while the schema defines the response contract.
A response from our Apple financial report test included data like:
{
"company": "Apple Inc.",
"reporting_period": "Q3 FY2026",
"period_end_date": "June 27, 2026",
"income_statement_three_months_ended": {
"total_net_sales": {
"current_period": 109417,
"prior_year_period": 94036
},
"net_income": {
"current_period": 29789,
"prior_year_period": 23434
},
"earnings_per_share_diluted": {
"current_period": 2.02,
"prior_year_period": 1.57
}
}
}Our full schema went much further than this. It extracted income statement figures, balance sheet data, cash flow, geographic revenue, product revenue, share repurchases, dividends, and other fields from a three-page Apple financial document. The extraction in that demo completed in roughly 10 seconds.
That is the useful difference between simply reading a PDF and turning it into data your application can work with. Once the result is structured JSON, you can store it in a database, compare reports across periods, feed it into a dashboard, or pass it into another application without building your own parser around the Markdown.
Spidra's scrape API is asynchronous. Submitting the request returns a jobId, which you use with GET /scrape/{jobId} until the job is complete. When you use a schema, the parsed JSON object is returned in result.content.
What if you do not know the PDF URL?
Many document workflows begin before scraping.
You might know that you need Apple's 2026 annual report, for example, without knowing where Apple has published the PDF.
Spidra Search supports a filetype: "pdf" filter for exactly this case:
curl -X POST https://api.spidra.io/api/search \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"query": "Apple annual report 2026",
"sources": ["web"],
"filetype": "pdf"
}'That restricts web results to PDFs.
If you also want the contents of the returned PDFs, add scrapeOptions:
curl -X POST https://api.spidra.io/api/search \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"query": "Apple annual report 2026",
"sources": ["web"],
"filetype": "pdf",
"limit": 5,
"scrapeOptions": {
"formats": ["markdown"]
}
}'Spidra then searches for the PDFs and fetches the document content before the Search job completes, with the Markdown added directly to each result. Search's scrapeOptions is intended for fetching page or document content as Markdown. If you want AI extraction or a JSON Schema, take the relevant URL and send it to the Scrape API instead.
That distinction keeps the workflow simple:
Search for the report
↓
Find relevant PDF results
↓
Read the documents as Markdown
↓
Send the document you need to /scrape
↓
Extract your JSON SchemaThis is useful for annual reports, research papers, filings, public reports, investor presentations, and other cases where you know what you are looking for but do not have the document URL yet.
A PDF is not always just text
One reason document scraping gets complicated is that two PDFs can look similar to a person while being very different underneath.
One may contain a proper text layer that can be extracted directly. Another may contain scanned pages that require OCR. A third may mix both.
Spidra handles scanned and image-only PDF pages as part of the document scraping flow, so you do not have to decide beforehand whether a file needs OCR.
The result is the same workflow from your side: provide the document URL and choose whether you want Markdown or structured extraction.
When should you use Markdown and when should you use JSON?
If you want the complete document content, Markdown is usually the better output.
PDF → MarkdownIf you already know which fields your application needs, use a schema.
PDF → JSON Schema → Structured JSONAnd if you still need to find the document:
Search → PDF → Scrape → Structured JSONFor example, a research application might keep the complete paper as Markdown. A financial data pipeline may only want revenue, net income, EPS, and balance sheet figures. A lead-enrichment workflow may want a fixed set of fields from hundreds of company documents.
The format should follow what you plan to do with the data afterward.
From documents on the web to usable data
PDF extraction often gets treated as a completely separate problem from web scraping, but many of the documents people want to process already live on the web.
Sometimes you have the URL and just need the contents. Sometimes you need a predictable JSON structure. Other times, you do not even know where the document is yet.
Spidra lets those steps live in the same workflow. Search can find the document, Scrape can read it, and a JSON Schema can define exactly what should come back.
You can explore Spidra Search, try document scraping from the Spidra dashboard, or use the Scrape API documentation to start building the workflow in code.
