Image-to-Text
Transformers
Safetensors
mistral3
text-generation
ocr
document-understanding
vision-language
pdf
tables
forms
Eval Results
๐ช๐บ Region: EU
Instructions to use lightonai/LightOnOCR-1B-1025 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use lightonai/LightOnOCR-1B-1025 with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "image-to-text" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # pip install "transformers<5.0.0" from transformers import pipeline pipe = pipeline("image-to-text", model="lightonai/LightOnOCR-1B-1025")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForSeq2SeqLM processor = AutoProcessor.from_pretrained("lightonai/LightOnOCR-1B-1025") model = AutoModelForSeq2SeqLM.from_pretrained("lightonai/LightOnOCR-1B-1025", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Without bounding box output
#11
by jasonni2 - opened
Congrats for this efficient model!
I'm using llama.cpp server for inference. When the document have digram, the ocr output only contains description of the diagram. Is this by design?
Thanks for trying the model!
yes the current version only indicates when there are images/charts/diagrams, we are working on adding bounding box information for each detected image natively by the model.
Stay tuned!