Opendataloader Pdf
haystack_integrations.components.converters.opendataloader_pdf.converter
OpenDataLoaderConverter
OpenDataLoader PDF converter component.
The component accepts PDF file paths and Haystack ByteStream objects, runs OpenDataLoader PDF extraction, and returns Haystack Document objects.
Parameters:
- output_format (
OutputFormat) – Output format generated by OpenDataLoader. Supported formats are markdown, text, json, and html. - convert_kwargs (
dict[str, Any] | None) – Additional arguments passed toopendataloader_pdf.convert. See the OpenDataLoader PDF Python options.
init
python
__init__(
*,
output_format: OutputFormat = "markdown",
convert_kwargs: dict[str, Any] | None = None
) -> None
Initialize the OpenDataLoader converter.
Parameters:
- output_format (
OutputFormat) – Format OpenDataLoader should produce. - convert_kwargs (
dict[str, Any] | None) – Additional arguments passed toopendataloader_pdf.convert. See the [OpenDataLoader PDF Python options] (https://opendataloader.org/docs/quick-start-python).
to_dict
Serialize the component.
Returns:
dict[str, Any]– Dictionary representation of the converter.
from_dict
Deserialize the component.
Parameters:
- data (
dict[str, Any]) – Serialized component dictionary.
Returns:
OpenDataLoaderConverter– Reconstructed OpenDataLoaderConverter.
run
python
run(
sources: list[str | Path | ByteStream],
meta: dict[str, Any] | list[dict[str, Any]] | None = None,
) -> dict[str, list[Document]]
Convert PDF sources into Haystack Documents.
Parameters:
- sources (
list[str | Path | ByteStream]) – PDF file paths or Haystack ByteStream objects. - meta (
dict[str, Any] | list[dict[str, Any]] | None) – Optional metadata attached to the generated Documents. A single dictionary is applied to every source. A list must contain one dictionary per source. ByteStream metadata is also preserved.
Returns:
dict[str, list[Document]]– Dictionary containing the converted Documents.