Convert a PDF into XML containing extracted text and information about its layout, such as font details and object positions.
IN THIS TUTORIAL
Step 1: Prepare Your Project
Install a currently supported Node.js version and obtain your API key from the PDF.co dashboard.
This example uses Node.js’s built-in fetch, so no additional packages are required.
Step 2: Create the Script
Save the following as app.js. Replace the API key placeholder and, if needed, the source PDF URL.
const fs = require("node:fs/promises");
const API_KEY = "YOUR_PDFCO_API_KEY";
const SOURCE_URL =
"https://pdfco-test-files.s3.us-west-2.amazonaws.com/pdf-to-xml/sample.pdf";
const OUTPUT_FILE = "result.xml";
async function main() {
const response = await fetch(
"https://api.pdf.co/v1/pdf/convert/to/xml",
{
method: "POST",
headers: {
"x-api-key": API_KEY,
"Content-Type": "application/json"
},
body: JSON.stringify({
url: SOURCE_URL,
name: OUTPUT_FILE,
pages: "",
inline: false,
async: false
})
}
);
if (!response.ok) {
throw new Error(`Conversion request failed: HTTP ${response.status}`);
}
const result = await response.json();
if (result.error || !result.url) {
throw new Error(result.message || "No output URL was returned.");
}
const download = await fetch(result.url);
if (!download.ok) {
throw new Error(`Download failed: HTTP ${download.status}`);
}
await fs.writeFile(
OUTPUT_FILE,
Buffer.from(await download.arrayBuffer())
);
console.log(`Saved ${OUTPUT_FILE}`);
}
main().catch((error) => {
console.error(error.message);
process.exitCode = 1;
});The script submits the PDF for conversion, checks the response, and downloads the XML to your project folder.
Leave pages empty to process every page, or use "0" for the first page. For a password-protected PDF, add a password property to the request body. See the PDF to XML API documentation for additional options and code samples.
Step 3: Run the Conversion
Open a terminal in your project folder and run:
node app.jsAfter a successful conversion, open result.xml in a text editor to inspect the extracted content.
This example uses synchronous processing for a small PDF. Large documents should use asynchronous processing and check the job status before downloading the result.
