Image-Text-to-Text
PaddleOCR
Safetensors
English
Chinese
multilingual
paddleocr_vl
ERNIE4.5
PaddlePaddle
image-to-text
ocr
document-parse
layout
table
formula
chart
seal
spotting
conversational
custom_code
Eval Results
Instructions to use PaddlePaddle/PaddleOCR-VL-1.5 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PaddleOCR
How to use PaddlePaddle/PaddleOCR-VL-1.5 with PaddleOCR:
# See https://www.paddleocr.ai/latest/version3.x/pipeline_usage/PaddleOCR-VL.html to installation from paddleocr import PaddleOCRVL pipeline = PaddleOCRVL(pipeline_version="v1.5") output = pipeline.predict("path/to/document_image.png") for res in output: res.print() res.save_to_json(save_path="output") res.save_to_markdown(save_path="output") - Notebooks
- Google Colab
- Kaggle
Personal Experience
#13
by fluxnad - opened
Personal thought. PaddleOCR-VL does a great job on text recognition. What I noticed in complex tables is a cell detection issue. When the table relies on alignment and spacing instead of clear cell borders, the model sometimes merges cells or assigns values to the wrong column. In my example, subtotal rows like “S/Total” lose the correct column alignment, and at times a whole column region gets treated as one cell when the structure is not clearly labeled and it droped the values of 2020
and this is the output of the ocr





