Files
docling/tests/data/groundtruth/docling_v2/example_09.html.md
krrome 94fcc46aa9 feat(html): Support formatting tags in HTML texts (#2111)
* add parsing for formatting tags in HTML backend

Signed-off-by: Roman Kayan BAZG <roman.kayan@bazg.admin.ch>

fix latest tests + wiki_duck result files.

Signed-off-by: Roman Kayan BAZG <roman.kayan@bazg.admin.ch>

* convert _collect_parent_format_tags to staticmethod

Signed-off-by: Roman Kayan BAZG <roman.kayan@bazg.admin.ch>

---------

Signed-off-by: Roman Kayan BAZG <roman.kayan@bazg.admin.ch>
2025-08-22 10:37:34 +02:00

728 B
Vendored

Introduction to parsing HTML files with Docling

Docling

Docling simplifies document processing, parsing diverse formats - including HTML - and providing seamless integrations with the gen AI ecosystem.

Supported file formats

Docling supports multiple file formats..

  • Advanced PDF understanding PDF
  • Microsoft Office DOCX DOCX
  • HTML files (with optional support for images) HTML

Three backends for handling HTML files

Docling has three backends for parsing HTML files:

  1. HTMLDocumentBackend Ignores images
  2. HTMLDocumentBackendImagesInline Extracts images inline
  3. HTMLDocumentBackendImagesReferenced Extracts images as references