Skip to main content
Version: 3.3

GotenbergFileConverter

A component that converts Office, OpenDocument, HTML, and Markdown files, and web pages to PDF using a Gotenberg server.

Most common position in a pipelineBefore a PDF converter (e.g. PyPDFToDocument) at the beginning of an indexing pipeline
Mandatory run variablessources: File paths, HTTP(S) URLs, or ByteStream objects
Output variablesoutput: A list of PDF ByteStream objects, one per source
API referenceGotenberg
GitHub linkhttps://github.com/deepset-ai/haystack-core-integrations/tree/main/integrations/gotenberg
Package namegotenberg-haystack

Overview​

GotenbergFileConverter sends files to a Gotenberg server and returns the converted PDFs. Gotenberg is a Docker-based API that wraps LibreOffice and Chromium, so you can run and scale document conversion as a separate service instead of installing LibreOffice next to your pipeline.

Each source is routed to a Gotenberg conversion endpoint on its own, so one batch can mix source types:

SourceGotenberg route
Local .md or .markdown fileChromium Markdown
Local .html, .htm, or .xhtml fileChromium HTML
Any other local file with a supported extension (.docx, .xlsx, .pptx, .odt, .rtf, ...)LibreOffice
String starting with http:// or https://Chromium URL. Gotenberg fetches the page and prints it to PDF.
ByteStreamRouted by its mime_type, which is required

Like LibreOfficeFileConverter, this component outputs ByteStream objects rather than Haystack Documents. Chain it with a PDF converter such as PyPDFToDocument to get Documents.

Formats that Haystack can already read, such as .docx, .pptx, .xlsx, .md, and .html, are usually best converted directly with their own converters, for example DOCXToDocument or MarkdownToDocument. Gotenberg only outputs PDF, so use this component for formats without a native Haystack converter, such as legacy Office files (.doc, .ppt, .xls), OpenDocument files (.odt, .ods, .odp), RTF, and Apple iWork files. It is also useful when you want a PDF rendering of a document or web page.

The output metadata keeps the local source path under file_path and any metadata of a source ByteStream. Metadata you pass through the meta run parameter takes precedence. HTML and Markdown sources can reference local assets such as images and stylesheets: pass them through the resources run parameter.

Use run_async to convert sources concurrently. concurrency_limit sets how many requests are sent to Gotenberg at once.

Running Gotenberg​

The component needs a running Gotenberg server. The easiest way to start one is with Docker:

shell
docker run --rm -p 3000:3000 gotenberg/gotenberg:8

By default, the component connects to http://localhost:3000. Use the url init parameter to point it to another server. See the Gotenberg documentation for more deployment options.

Usage​

Install the Gotenberg integration:

shell
pip install gotenberg-haystack

On its own​

python
from pathlib import Path

from haystack_integrations.components.converters.gotenberg import (
GotenbergFileConverter,
)

converter = GotenbergFileConverter(url="http://localhost:3000")
result = converter.run(
sources=[Path("report.docx"), Path("notes.md"), "https://haystack.deepset.ai"],
)
pdfs = result["output"]

To convert an HTML ByteStream together with the stylesheet it links to:

python
from pathlib import Path

from haystack.dataclasses import ByteStream
from haystack_integrations.components.converters.gotenberg import (
GotenbergFileConverter,
)

converter = GotenbergFileConverter()
html = ByteStream.from_file_path(Path("invoice.html"), mime_type="text/html")
result = converter.run(
sources=[html],
meta={"customer": "ACME"},
resources=[Path("style.css")],
)

In a pipeline​

The pipeline below converts a legacy Word document and an OpenDocument text file to PDF, extracts their text, splits it, and writes the Documents to a Document Store.

PyPDFToDocument and sentence splitting in DocumentSplitter need extra dependencies:

shell
pip install gotenberg-haystack pypdf nltk
python
from pathlib import Path

from haystack import Pipeline
from haystack.components.converters import PyPDFToDocument
from haystack.components.preprocessors import DocumentSplitter
from haystack.components.writers import DocumentWriter
from haystack.document_stores.in_memory import InMemoryDocumentStore
from haystack_integrations.components.converters.gotenberg import (
GotenbergFileConverter,
)

document_store = InMemoryDocumentStore()

pipeline = Pipeline()
pipeline.add_component("gotenberg", GotenbergFileConverter())
pipeline.add_component("pdf_converter", PyPDFToDocument())
pipeline.add_component(
"splitter",
DocumentSplitter(split_by="sentence", split_length=5),
)
pipeline.add_component("writer", DocumentWriter(document_store=document_store))
pipeline.connect("gotenberg.output", "pdf_converter.sources")
pipeline.connect("pdf_converter", "splitter")
pipeline.connect("splitter", "writer")

pipeline.run({"gotenberg": {"sources": [Path("report.doc"), Path("notes.odt")]}})