docling/tests/data/groundtruth/docling_v2/unit_test_01.html.itxt
Alexander Vaagan 733360c7b2 A new HTML backend that handles styled html (ignors it) as well as images.
- Updated unit tests
- Added documentation (Example notebook)

Note: MyPy fails.
Seems to be a known issue with BeautifulSoup:
https://github.com/python/typeshed/pull/13604

Signed-off-by: Alexander Vaagan <alexander.vaagan@gmail.com>
Signed-off-by: vaaale <2428222+vaaale@users.noreply.github.com>
2025-05-24 22:29:22 +02:00

8 lines
370 B
Plaintext

item-0 at level 0: unspecified: group _root_
item-1 at level 1: title: Title
item-2 at level 1: section_header: section-1
item-3 at level 1: section_header: section-1.1
item-4 at level 1: section_header: section-2
item-5 at level 1: section_header: section-2.0.1
item-6 at level 1: section_header: section-2.2
item-7 at level 1: section_header: section-2.3